Network Model Data Processing, Data Display Method, Device and Storage Medium
By using a multi-level cached queue mechanism in the intelligent learning platform, the problem of inefficient model training is solved, and faster data access and training process is achieved.
Patent Information
- Application Number
- CN202110494425.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-07
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-05-07
AI Technical Summary
The intelligent learning platform needs to download and parse a large amount of data during model training, resulting in inefficient training.
By setting the first queue and the second queue in the data cache, the multi-level caching mechanism is used to find and read the training data set, and direct access to the disk is reduced.
Improves the efficiency of model training and reduces the time overhead of downloading, parsing and transmission.
Smart Images

Figure CN113761004B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method, apparatus, computer device, and storage medium for network model data processing and data display. Background Art
[0002] With the development of artificial intelligence technology, intelligent learning platform technology has emerged. An intelligent learning platform is a one-stop machine learning ecosystem service platform with powerful computing capabilities. It can combine various data sources, components, algorithms, models, and evaluation modules, enabling algorithm engineers and data scientists to conveniently perform model training, evaluation, and prediction on it. It supports massive data storage, annotation, inference, training, and distribution. During the process of intelligent learning and processing, it is necessary to display massive data. Currently, when an intelligent learning platform trains a model, it usually downloads and parses the training data set to be used from the medium storing the data, and then performs model training. However, the intelligent learning platform often uses data to train multiple different models, and each time a model is trained, it will trigger time-consuming operations such as full-scale download, parsing, and transmission of the data set, resulting in a decrease in the efficiency of model training. Summary of the Invention
[0003] Based on this, in view of the above technical problems, it is necessary to provide a method, apparatus, computer device, and storage medium for network model data processing and data display that can improve the efficiency of model training.
[0004] A method for network model data processing, the method includes:
[0005] Receiving a training data acquisition request, where the training data acquisition request carries an identifier of the training data set to be acquired;
[0006] When the type of the training data set to be acquired corresponding to the identifier of the training data set to be acquired is a text type, searching for the corresponding target text training data set identifier in the first queue; the first queue stores each first text training data set identifier and the corresponding first text training data set;
[0007] When the corresponding target text training data set identifier is not found in the first queue, searching for the corresponding target text training data set identifier of the training data set to be acquired in the second queue; the second queue stores each second text training data set identifier, the corresponding historical access times, and the corresponding disk storage location, and both the first queue and the second queue are located in the data cache;
[0008] When the corresponding target text training data set identifier is found in the second queue, obtaining the corresponding target disk storage location of the found target text training data set identifier;
[0009] Read the target text training dataset corresponding to the target text training dataset identifier from the corresponding disk based on the target disk storage location, and update the historical access count corresponding to the target text training dataset identifier in the second queue. The second text training dataset corresponding to the second text training dataset identifier in the second queue is used to be written into the first queue when the corresponding historical access count meets the cache condition;
[0010] Input the target text training dataset into the network model for training. The network model is used to process the input data according to the model task type.
[0011] A network model data processing device, the device includes:
[0012] A request receiving module, configured to receive a training data acquisition request, and the training data acquisition request carries an identifier of the training dataset to be acquired;
[0013] A first search module, configured to, when the type of the training dataset to be acquired corresponding to the identifier of the training dataset to be acquired is a text type, search for the corresponding target text training dataset identifier in the first queue; the first queue stores various first text training dataset identifiers and the corresponding first text training datasets;
[0014] A second search module, configured to, when the corresponding target text training dataset identifier is not found in the first queue, search for the corresponding target text training dataset identifier of the training dataset to be acquired in the second queue; the second queue stores various second text training dataset identifiers, the corresponding historical access counts, and the corresponding disk storage locations, and both the first queue and the second queue are located in the data cache;
[0015] A location acquisition module, configured to, when the corresponding target text training dataset identifier is found in the second queue, acquire the target disk storage location corresponding to the found target text training dataset identifier;
[0016] A data reading module, configured to read the target text training dataset corresponding to the target text training dataset identifier from the corresponding disk based on the target disk storage location, and update the historical access count corresponding to the target text training dataset identifier in the second queue. The second text training dataset corresponding to the second text training dataset identifier in the second queue is used to be written into the first queue when the corresponding historical access count meets the cache condition;
[0017] A training module, configured to input the target text training dataset into the network model for training. The network model is used to process the input data according to the model task type.
[0018] A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented:
[0019] Receive a training data acquisition request, where the training data acquisition request carries an identifier of a training data set to be acquired;
[0020] When the type of the training data set to be acquired corresponding to the identifier of the training data set to be acquired is a text type, search for the corresponding target text training data set identifier in the first queue; the first queue stores various first text training data set identifiers and corresponding first text training data sets;
[0021] When the corresponding target text training data set identifier is not found in the first queue, search for the corresponding target text training data set identifier in the second queue; the second queue stores various second text training data set identifiers, corresponding historical access times, and corresponding disk storage locations. Both the first queue and the second queue are located in the data cache;
[0022] When the corresponding target text training data set identifier is found in the second queue, obtain the corresponding target disk storage location of the found target text training data set identifier;
[0023] Based on the target disk storage location, read the target text training data set corresponding to the target text training data set identifier from the corresponding disk, update the historical access times corresponding to the target text training data set identifier in the second queue, and the second text training data sets corresponding to the second text training data set identifiers in the second queue are used to be written into the first queue when the corresponding historical access times meet the cache conditions;
[0024] Input the target text training data set into a network model for training, and the network model is used to process the input data according to the model task type.
[0025] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0026] Receive a training data acquisition request, where the training data acquisition request carries an identifier of a training data set to be acquired;
[0027] When the type of the training data set to be acquired corresponding to the identifier of the training data set to be acquired is a text type, search for the corresponding target text training data set identifier in the first queue; the first queue stores various first text training data set identifiers and corresponding first text training data sets;
[0028] When the corresponding target text training dataset identifier cannot be found in the first queue, search for the target text training dataset identifier corresponding to the training dataset identifier to be obtained in the second queue; the second queue stores various second text training dataset identifiers, corresponding historical access times, and corresponding disk storage locations, and both the first queue and the second queue are located in the data cache;
[0029] When the corresponding target text training dataset identifier is found in the second queue, obtain the target disk storage location corresponding to the found target text training dataset identifier;
[0030] Read the target text training dataset corresponding to the target text training dataset identifier from the corresponding disk based on the target disk storage location, update the historical access times corresponding to the target text training dataset identifier in the second queue, and the second text training dataset corresponding to the second text training dataset identifier in the second queue is used to be written into the first queue when the corresponding historical access times meet the cache conditions;
[0031] Input the target text training dataset into the network model for training, and the network model is used to process the input data according to the model task type.
[0032] The above network model data processing method, device, computer device, and storage medium receive a training data acquisition request, and the training data acquisition request carries the training dataset identifier to be obtained. When the type of the training dataset to be obtained corresponding to the training dataset identifier to be obtained is a text type, first search for the corresponding target text training dataset identifier in the first queue. When the corresponding target text training dataset identifier cannot be found in the first queue, search for the target text training dataset identifier corresponding to the training dataset identifier to be obtained in the second queue, and both the first queue and the second queue are located in the data cache. When the corresponding target text training dataset identifier is found in the second queue, obtain the target disk storage location corresponding to the found target text training dataset identifier, read the target text training dataset corresponding to the target text training dataset identifier from the corresponding disk based on the target disk storage location, and then input the target text training dataset into the network model for training. That is, through the first queue and the second queue, multi-level caching is used for access when obtaining training data, making full use of the cache and the disk to reduce the time consumed by downloading, parsing, and transmission, etc., avoiding the server from downloading the data to be displayed from the storage medium, so that training data can be quickly obtained, and further improving the efficiency of the training model.
[0033] A data display method, the method includes:
[0034] Receive a data display instruction sent through the data display page, and send a data acquisition request to the server based on the data display instruction. The data acquisition request carries the dataset identifier to be displayed,
[0035] The server is used to, when the type of the dataset to be displayed corresponding to the dataset identifier to be displayed is a text type, search for the corresponding target text dataset identifier in the first queue based on the dataset identifier to be displayed. The first queue stores each first text dataset identifier and the corresponding first text dataset. When the corresponding target text dataset identifier is not found in the first queue, search for the corresponding target text dataset identifier of the dataset to be displayed in the second queue. The second queue stores each second text dataset identifier, the corresponding historical access times, and the corresponding disk storage location. Both the first queue and the second queue are located in the data cache. When the corresponding target text dataset identifier is found in the second queue, obtain the target disk storage location corresponding to the found target text dataset identifier, read the target text dataset corresponding to the target text dataset identifier from the corresponding disk based on the target disk storage location, update the historical access times corresponding to the target text dataset identifier in the second queue. The second text dataset corresponding to the second text dataset identifier in the second queue is used to be written into the first queue when the corresponding historical access times meet the cache condition, and return the target text dataset to the terminal;
[0036] Obtain the target text dataset returned by the server, and display the target text dataset through the data display page.
[0037] In one embodiment, the data acquisition request carries a dataset identifier to be displayed and a requester identifier;
[0038] The sending a data acquisition request to the server based on the data display instruction includes:
[0039] Send a target data acquisition request to the server based on the data display instruction. The server parses the target data acquisition request to obtain the requester identifier, and matches the requester identifier with a preset data reading permission list. When there is a matching requester identifier, the server responds to the data synchronization request, generates an asynchronous task identifier, records the asynchronous task state corresponding to the asynchronous task identifier as unfinished, and returns the asynchronous task identifier to the terminal.
[0040] In one embodiment, the method further includes:
[0041] Receive a data display instruction sent through the data display page, and send a data acquisition request to the server based on the data display instruction. The data acquisition request carries a dataset identifier to be displayed;
[0042] The server is configured to, when the type of the dataset to be displayed corresponding to the dataset identifier to be displayed is of text type, search for the corresponding target text dataset identifier in the first queue based on the dataset identifier to be displayed. When the corresponding target text dataset identifier is found in the first queue, obtain the target cache storage location corresponding to the target text dataset identifier, read the target text dataset corresponding to the target text dataset identifier from the corresponding cache based on the target cache storage location; obtain the current time point, update the first access time point corresponding to the target text dataset identifier based on the current time point, and return the target text dataset corresponding to the dataset identifier to be displayed to the terminal;
[0043] Obtain the target text dataset returned by the server, and display the target text dataset through the data display page.
[0044] A data display device, the device includes:
[0045] A request sending module, configured to receive a data display instruction sent through the data display page, and send a data acquisition request to the server based on the data display instruction. The data acquisition request carries the dataset identifier to be displayed. The server is configured to, when the type of the dataset to be displayed corresponding to the dataset identifier to be displayed is of text type, search for the corresponding target text dataset identifier in the first queue based on the dataset identifier to be displayed. The first queue stores each first text dataset identifier and the corresponding first text dataset. When the corresponding target text dataset identifier is not found in the first queue, search for the target text dataset identifier corresponding to the dataset identifier to be displayed in the second queue. The second queue stores each second text dataset identifier, the corresponding historical access times, and the corresponding disk storage location. Both the first queue and the second queue are located in the data cache. When the corresponding target text dataset identifier is found in the second queue, obtain the target disk storage location corresponding to the found target text dataset identifier, read the target text dataset corresponding to the target text dataset identifier from the corresponding disk based on the target disk storage location, update the historical access times corresponding to the target text dataset identifier in the second queue. The second text dataset corresponding to the second text dataset identifier in the second queue is configured to be written into the first queue when the corresponding historical access times meet the cache condition, and return the target text dataset to the terminal;
[0046] A text display module, configured to obtain the target text dataset returned by the server, and display the target text dataset through the data display page.
[0047] A computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0048] Receive the data display instruction sent through the data display page, and send a data acquisition request to the server based on the data display instruction. The data acquisition request carries the identifier of the data set to be displayed.
[0049] The server is used to, when the type of the data set to be displayed corresponding to the identifier of the data set to be displayed is a text type, search for the corresponding target text data set identifier in the first queue based on the identifier of the data set to be displayed. The first queue stores each first text data set identifier and the corresponding first text data set. When the corresponding target text data set identifier is not found in the first queue, search for the corresponding target text data set identifier of the data set to be displayed in the second queue. The second queue stores each second text data set identifier, the corresponding historical access times, and the corresponding disk storage location. Both the first queue and the second queue are located in the data cache. When the corresponding target text data set identifier is found in the second queue, obtain the corresponding target disk storage location of the found target text data set identifier, read the target text data set corresponding to the target text data set identifier from the corresponding disk based on the target disk storage location, update the historical access times corresponding to the target text data set identifier in the second queue. The second text data set corresponding to the second text data set identifier in the second queue is used to be written into the first queue when the corresponding historical access times meet the cache condition, and return the target text data set to the terminal.
[0050] Obtain the target text data set returned by the server, and display the target text data set through the data display page.
[0051] A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0052] Receive the data display instruction sent through the data display page, and send a data acquisition request to the server based on the data display instruction. The data acquisition request carries the identifier of the data set to be displayed.
[0053] The server is used to, when the type of the dataset to be displayed corresponding to the dataset identifier to be displayed is a text type, search for the corresponding target text dataset identifier in the first queue. The first queue stores each first text dataset identifier and the corresponding first text dataset. When the corresponding target text dataset identifier is not found in the first queue, search for the target text dataset identifier corresponding to the dataset identifier to be displayed in the second queue. The second queue stores each second text dataset identifier, the corresponding historical access times, and the corresponding disk storage location. Both the first queue and the second queue are located in the data cache. When the corresponding target text dataset identifier is found in the second queue, obtain the target disk storage location corresponding to the found target text dataset identifier, read the target text dataset corresponding to the target text dataset identifier from the corresponding disk based on the target disk storage location, update the historical access times corresponding to the target text dataset identifier in the second queue. The second text dataset corresponding to the second text dataset identifier in the second queue is used to be written into the first queue when the corresponding historical access times meet the cache condition, and return the target text dataset to the terminal;
[0054] Obtain the target text dataset returned by the server, and display the target text dataset through the data display page.
[0055] For the above data display method, device, computer device, and storage medium, the terminal receives a data display instruction sent through the data display page, and sends a data acquisition request to the server based on the data display instruction. The data acquisition request carries the dataset identifier to be displayed. Then, the server is used to, when the type of the dataset to be displayed corresponding to the dataset identifier to be displayed is a text type, search for the corresponding target text dataset identifier in the first queue. When the corresponding target text dataset identifier is not found in the first queue, search for the target text dataset identifier corresponding to the dataset identifier to be displayed in the second queue. Both the first queue and the second queue are located in the data cache. When the corresponding target text dataset identifier is found in the second queue, obtain the target disk storage location corresponding to the found target text dataset identifier, read the target text dataset corresponding to the target text dataset identifier from the corresponding disk based on the target disk storage location, and return the target text dataset to the terminal; Obtain the target text dataset returned by the server, and display the target text dataset through the data display page. That is, through the first queue and the second queue, multi-level caching is used for access when obtaining display data, making full use of the cache and the disk to reduce the time consumed by downloading, parsing, and transmission, etc., avoiding the server from downloading the data to be displayed from the storage medium, so that the data to be displayed can be quickly obtained, and the display speed of the data is improved. Description of the Drawings
[0056] Figure 1It is an application environment diagram of the network model data processing method in an embodiment;
[0057] Figure 2 It is a schematic flowchart of the network model data processing method in an embodiment;
[0058] Figure 3 It is a schematic flowchart of the network model data processing method in another embodiment;
[0059] Figure 4 It is a schematic flowchart of the network model data processing method in yet another embodiment;
[0060] Figure 5 It is a schematic flowchart of the network model data processing method in still another embodiment;
[0061] Figure 6 It is a schematic flowchart of the process of writing to the second queue in an embodiment;
[0062] Figure 7 It is a schematic flowchart of the update process of the second queue in a specific embodiment;
[0063] Figure 8 It is a schematic diagram of the interaction between queues in a specific embodiment;
[0064] Figure 9 It is a schematic flowchart of the update process of writing to the first queue in an embodiment;
[0065] Figure 10 It is a schematic flowchart of the use of the first queue in a specific embodiment;
[0066] Figure 11 It is a schematic diagram of the first queue performing data processing in a specific embodiment;
[0067] Figure 12 It is a schematic flowchart of the data display method in an embodiment;
[0068] Figure 13 It is a schematic flowchart of the data display method in another embodiment;
[0069] Figure 14 It is a schematic flowchart of the data display method in yet another embodiment;
[0070] Figure 15 It is a specific framework diagram of the text data set display in a specific embodiment;
[0071] Figure 16 It is a schematic flowchart of the data display method in still another embodiment;
[0072] Figure 17Schematic diagram of the specific framework shown in the picture data set in a specific embodiment;
[0073] Figure 18 Flow chart of the network model data processing method in a specific embodiment;
[0074] Figure 19 Schematic diagram of the framework of the data center system in a specific embodiment;
[0075] Figure 20 For Figure 19 Flow chart of data display in a specific embodiment;
[0076] Figure 21 For Figure 19 Schematic diagram of a partial text data set shown in a specific embodiment;
[0077] Figure 22 For Figure 19 Flow chart of a partial picture data set shown in a specific embodiment;
[0078] Figure 23 Structural block diagram of the network model data processing device in an embodiment;
[0079] Figure 24 Structural block diagram of the data display device in an embodiment;
[0080] Figure 25 Internal structure diagram of a computer device in an embodiment;
[0081] Figure 26 Internal structure diagram of a computer device in another embodiment. Detailed implementation manners
[0082] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0083] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement, and further performs image processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0084] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, knowledge graph, etc.
[0085] The solution provided by the embodiments of this application involves technologies such as computer vision technology and natural language processing in artificial intelligence, and is specifically described through the following embodiments:
[0086] The network model data processing method provided by this application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The training data is stored in the database 106, and the server 104 can obtain the training data from the database 106. The terminal 102 receives a training instruction through the intelligent learning platform and sends a model training instruction to the server 104. The server 104 receives the model training instruction. According to the model training instruction, the server 104 receives a training data acquisition request, and the training data acquisition request carries an identifier of the training data set to be acquired. When the type of the training data set to be acquired corresponding to the identifier of the training data set to be acquired is a text type, the corresponding target text training data set identifier is searched in the first queue based on the identifier of the training data set to be acquired. The first queue stores each first text training data set identifier and the corresponding first text training data set. When the server 104 fails to find the corresponding target text training data set identifier in the first queue, the corresponding target text training data set identifier of the identifier of the training data set to be acquired is searched in the second queue. The second queue stores each second text training data set identifier, the corresponding historical access times, and the corresponding disk storage location. Both the first queue and the second queue are located in the data cache. When the server 104 finds the corresponding target text training data set identifier in the second queue, the corresponding target disk storage location of the found target text training data set identifier is obtained. Based on the target disk storage location, the target text training data set corresponding to the target text training data set identifier is read from the corresponding disk, and the historical access times corresponding to the target text training data set identifier in the second queue are updated. The second text training data set corresponding to the second text training data set identifier in the second queue is used to be written into the first queue when the corresponding historical access times meet the cache condition. The target text training data set is input into the network model for training, and the network model is used to process the input data according to the model task type. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.
[0087] In one embodiment, as Figure 2 shown, a network model data processing method is provided, and this method is applied to Figure 1Taking the server in as an example for illustration, it can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The method includes the following steps:
[0088] Step 202: Receive a training data acquisition request, where the training data acquisition request carries an identifier of the training data set to be acquired.
[0089] Among them, the training data acquisition request is used to request the acquisition of training data. Training data refers to the data used when training an artificial intelligence model, which can be text data or image data. An artificial intelligence model refers to a model established through artificial intelligence algorithms, and the artificial intelligence algorithms include supervised learning algorithms and unsupervised learning algorithms, etc. Among them, the supervised learning algorithms can be decision tree algorithms, neural network algorithms, linear regression algorithms, etc., and the unsupervised learning algorithms can be clustering algorithms, adversarial neural network algorithms, etc. The identifier of the training data set to be acquired is used to uniquely identify the training data set to be acquired, and the training data set to be acquired refers to the training data set that needs to be acquired.
[0090] Specifically, the server can receive a training data acquisition request, and the training data acquisition request carries an identifier of the training data set to be acquired. Among them, it can be that after the terminal sends a training data acquisition request to the server, the server receives the training data acquisition request, or it can be that the server receives the training data acquisition request when executing a model training instruction.
[0091] Step 204: When the type of the training data set to be acquired corresponding to the identifier of the training data set to be acquired is of text type, search for the corresponding target text training data set identifier in the first queue. The first queue stores various first text training data set identifiers and the corresponding first text training data sets.
[0092] Among them, the type of the training data set to be acquired is used to characterize the type of the training data set, including text type and image type. The text type refers to the type stored in text format, and the image type refers to the type stored in image format. The first queue stores various first text training data set identifiers and the corresponding first text training data sets. The first text training data set identifier refers to the text training data set identifier cached in the first queue. The text training data set identifier is used to uniquely identify the text training data set. The first text training data set refers to the text training data set cached in the first queue. That is, the first queue caches text training data sets, and can cache all text training data sets, or can cache some text training data in the text training data sets. The target text training data set identifier refers to the identifier of the text training data set to be acquired.
[0093] Specifically, the server first determines the type of training data to be obtained, that is, it obtains the format of the training data set corresponding to the identifier of the training data set to be obtained, and determines the type of the training data set to be obtained according to the format of the training data set. When the format of the training data set is text format, that is, the type of the training data set to be obtained corresponding to the identifier of the training data set to be obtained is text type. At this time, search in the first queue using the identifier of the training data set to be obtained. When a consistent identifier is found, the consistent identifier is the target text training data set identifier corresponding to the identifier of the training data set to be obtained in the first queue. Among them, each first text training data set identifier and the corresponding first text training data set are stored in the first queue, and the number of first text training data set identifiers stored in the first queue can be set according to requirements.
[0094] Step 206, when the corresponding target text training data set identifier is not found in the first queue, search for the target text training data set identifier corresponding to the identifier of the training data set to be obtained in the second queue; each second text training data set identifier, the corresponding historical access times, and the corresponding disk storage location are stored in the second queue, and both the first queue and the second queue are located in the data cache.
[0095] Among them, each second text training data set identifier, the corresponding historical access times, and the corresponding disk storage location are stored in the second queue. Among them, the second text training data set identifier refers to the text training data set identifier cached in the second queue, the historical access times refers to the number of times the text training data set corresponding to the second text training data set identifier in the second queue has been historically accessed, and each time it is accessed, the historical access times will increase by one. The disk storage location refers to the specific storage location of the text training data corresponding to the second text training data identifier in the disk. Both the first queue and the second queue are queues in the cache.
[0096] Specifically, if the server does not find the corresponding target text training data set identifier in the first queue, it means that the training data to be obtained is not stored in the first queue. At this time, the server uses the identifier of the training data set to be obtained to match each second text training data set identifier stored in the second queue, that is, searches for the target text training data set identifier corresponding to the identifier of the training data set to be obtained in the second queue. Among them, each second text training data set identifier, the corresponding historical access times, and the corresponding disk storage location are associated and stored in the second queue, and both the first queue and the second queue are located in the data cache. That is, both the first queue and the second queue are stored in the data cache, and the server can directly search and use them from the data cache, which can improve efficiency.
[0097] Step 208, when the corresponding target text training dataset identifier is found in the second queue, obtain the target disk storage location corresponding to the found target text training dataset identifier.
[0098] The target disk storage location refers to the storage location of the training data to be obtained on the disk.
[0099] Specifically, when the server matches a consistent training dataset identifier to be obtained in the second queue, it means that the corresponding target text training dataset identifier is found in the second queue. At this time, obtain the disk storage location corresponding to the found target text training dataset identifier according to the association relationship in the second queue, and this disk storage location is the target disk storage location.
[0100] Step 210, based on the target disk storage location, read the target text training dataset corresponding to the target text training dataset identifier from the corresponding disk, update the historical access times corresponding to the target text training dataset identifier in the second queue, and the second text training dataset corresponding to the second text training dataset identifier in the second queue is used to be written into the first queue when the corresponding historical access times meet the cache condition.
[0101] The target text training dataset refers to the training data to be obtained, that is, the training data corresponding to the training dataset identifier to be obtained. The cache condition refers to the condition set in advance for storing the text training dataset corresponding to the second text training dataset identifier in the second queue into the first queue, which may include that the historical access times meet the preset number threshold.
[0102] Specifically, the server reads the target text training dataset corresponding to the target text training dataset identifier from the corresponding disk according to the target disk storage location, and at the same time, the server updates the historical access times corresponding to the target text training dataset identifier in the second queue, that is, the server increases the historical access times corresponding to the target text training dataset identifier in the second queue. Among them, the second text training dataset corresponding to the second text training dataset identifier in the second queue is used to be written into the first queue when the corresponding historical access times meet the cache condition.
[0103] Step 212, input the target text training dataset into the network model for training, and the network model is used to process the input data according to the model task type.
[0104] Among them, the network model refers to an artificial intelligence network model trained using training data. The network model is used to process input data according to the model task type, and the model task type is determined according to needs. For example, if the model task type is classification, the network model is a model used to classify input data. Another example is that if the model task type is recognition, the network model is a model used to recognize input data. Still another example is that if the model task type is prediction, the network model is a model that uses input data for prediction.
[0105] Specifically, when the server obtains the target text training data set, it uses the target text training data set to train the network model to obtain the trained network model, and then uses the trained network model to process the input data.
[0106] For the above network model data processing method, by receiving a training data acquisition request, the training data acquisition request carries an identifier of the training data set to be acquired. When the type of the training data set to be acquired corresponding to the identifier of the training data set to be acquired is of text type, first search for the corresponding target text training data set identifier in the first queue. When the corresponding target text training data set identifier is not found in the first queue, search for the target text training data set identifier corresponding to the identifier of the training data set to be acquired in the second queue. Both the first queue and the second queue are located in the data cache. When the corresponding target text training data set identifier is found in the second queue, obtain the target disk storage location corresponding to the found target text training data set identifier, and read the target text training data set corresponding to the target text training data set identifier from the corresponding disk based on the target disk storage location, and then input the target text training data set into the network model for training. That is, through the first queue and the second queue, multi-level caching is used for access when obtaining training data, making full use of the cache and the disk to reduce the time consumed by downloading, parsing, and transmission, etc., avoiding the server from downloading the data to be displayed from the storage medium, so that training data can be quickly obtained, and further improving the efficiency of the training model.
[0107] In one embodiment, the first queue stores various first text training data set identifiers, corresponding first text training data sets, corresponding cache storage locations, and corresponding first access time points;
[0108] As Figure 3 shown, after step 204, that is, after searching for the corresponding target text training data set identifier in the first queue based on the identifier of the training data set to be acquired, it further includes:
[0109] Step 302, when the corresponding target text training dataset identifier is found in the first queue, obtain the target cache storage location corresponding to the target text training dataset identifier, and read the target text training dataset corresponding to the target text training dataset identifier from the corresponding cache based on the target cache storage location.
[0110] Among them, the cache storage location refers to the specific storage location of the text training dataset corresponding to the first text training dataset identifier in the cache in the first queue. The first access time point refers to the time point when the text training dataset corresponding to the first text training dataset identifier in the first queue was last accessed. The target cache storage location refers to the cache storage location corresponding to the target text training dataset identifier, that is, the cache storage location corresponding to the training dataset identifier to be obtained. In one embodiment, the first text training dataset and the corresponding cache storage location in the server can be stored in the first queue in the format of key-value pairs.
[0111] Specifically, if the server directly finds the corresponding target text training dataset identifier in the first queue, obtain the target cache storage location corresponding to the target text training dataset identifier according to the association relationship in the first queue, and read the target text training dataset corresponding to the target text training dataset identifier from the corresponding cache using the target cache storage location.
[0112] Step 304, obtain the current time point, and update the first access time point corresponding to the target text training dataset identifier based on the current time point.
[0113] Among them, the current time point refers to the time point when the target text training dataset is currently read.
[0114] Specifically, the server obtains the current time point and uses the current time point to replace the first access time point corresponding to the target text training dataset identifier in the first queue.
[0115] In the above embodiment, by storing the association relationship between the first text training dataset identifier and the cache storage location in the first queue, the server can quickly read the target text training dataset corresponding to the target text training dataset identifier from the target cache storage location, that is, it can achieve reading in constant time and improve the efficiency.
[0116] In one embodiment, the training data acquisition request also carries the data volume of the training dataset to be obtained;
[0117] Such as Figure 4As shown, in step 302, when the corresponding target text training dataset identifier is found in the first queue, the target cache storage location corresponding to the target text training dataset identifier is obtained, and the target text training dataset corresponding to the target text training dataset identifier is read from the corresponding cache based on the target cache storage location, including:
[0118] Step 402, when the corresponding target text training dataset identifier is found in the first queue, obtain the cache data volume of the target text training dataset corresponding to the target text training dataset identifier in the first queue.
[0119] Among them, the training dataset data volume to be obtained refers to the size of the training dataset to be obtained, and this size can be set according to requirements. The cache data volume refers to the data volume of the text training dataset cached in the first queue, which can be stored in full or part of the data can be stored.
[0120] Specifically, when the server receives a training data acquisition request and parses the training data acquisition request, the training dataset data volume to be obtained is also obtained. At this time, if the server finds the corresponding target text training dataset identifier in the first queue, it obtains the cache data volume of the target text training dataset corresponding to the target text training dataset identifier in the first queue.
[0121] Step 404, when the cache data volume exceeds the training dataset data volume to be obtained, obtain the target cache storage location corresponding to the target text training dataset identifier, and read the target text training dataset with the training dataset data volume to be obtained from the corresponding cache based on the target cache storage location.
[0122] Specifically, the server compares the size of the cache data volume and the training dataset data volume to be obtained. When the cache data volume exceeds the training dataset data volume to be obtained, it means that all the training data to be obtained has been cached in the first queue. At this time, the target cache storage location corresponding to the target text training dataset identifier is obtained from the first queue, and then the target text training dataset with the training dataset data volume to be obtained is read from the corresponding cache using the target cache storage location.
[0123] In one embodiment, after step 402, that is, after obtaining the cache data volume of the target text training dataset corresponding to the target text training dataset identifier in the first queue when the corresponding target text training dataset identifier is found in the first queue, the following steps are further included:
[0124] When the cache data volume does not exceed the training dataset data volume to be obtained, obtain the corresponding target disk storage location according to the target text training dataset identifier, and read the target text training dataset with the training dataset data volume to be obtained from the corresponding disk based on the target disk storage location.
[0125] Among them, the target disk storage location refers to the disk storage location corresponding to the target text training dataset.
[0126] Specifically, the server compares the amount of cached data with the amount of training dataset data to be obtained. When the amount of cached data does not exceed the amount of training dataset data to be obtained, it indicates that the target text training dataset cached in the first queue cannot meet the needs. At this time, the server directly searches for the corresponding target disk storage location according to the identifier of the target text training dataset, and reads the target text training dataset with the amount of training dataset data to be obtained from the corresponding disk using the target disk storage location. Among them, the server pre-stores the identifiers of each training dataset and the corresponding full-scale training data.
[0127] In a specific embodiment, the training data acquisition request may carry the number range of the text training data to be obtained. For example, it is necessary to obtain the text training data corresponding to numbers 1 to 100 in the target text training dataset. At this time, the server obtains the numbers of the text training datasets in the target text training dataset cached in the first queue, for example, numbers 1 to 200. Then the server compares the maximum value of the number range of the text training data to be obtained with the maximum value of the codes of the text training datasets in the cached target text training dataset. When the maximum value of the codes of the text training datasets in the cached target text training dataset exceeds the maximum value of the number range of the text training data to be obtained, that is, 200 exceeds 100, the text training data corresponding to numbers 1 to 100 is read from the target cache storage location corresponding to the target text training dataset identifier in the first queue to obtain the target text training dataset. When it is necessary to obtain the text training data corresponding to numbers 1 to 300 in the target text training dataset, at this time, the maximum value of the number range of the text training data to be obtained exceeds the maximum value of the codes of the text training datasets in the cached target text training dataset, that is, 300 exceeds 200. Then the server directly obtains the corresponding disk storage location according to the text training data identifier, and reads the target text training dataset with the amount of training dataset data to be obtained from the corresponding disk using the disk storage location.
[0128] In the above embodiment, by obtaining the amount of training dataset data to be obtained, comparing the amount of training dataset data to be obtained with the amount of cached data, and then determining whether to read the target text training data from the first queue, that is, partial text training data can be cached in the first queue, which can reduce the pressure on the cache.
[0129] In one embodiment, as Figure 5 shown, after step 206, that is, after searching for the target text training dataset identifier corresponding to the training dataset identifier to be obtained in the second queue, it further includes:
[0130] Step 502, when the target text training dataset identifier corresponding to the to-be-obtained training dataset identifier is not found in the second queue, obtain the download address of the to-be-obtained training dataset based on the to-be-obtained training dataset identifier.
[0131] The training dataset download address refers to the download address when the training dataset is downloaded from the storage medium. The storage medium is a medium for storing data, which can be a distributed server, a blockchain, etc.
[0132] Specifically, when the server does not find the target text training dataset identifier corresponding to the to-be-obtained training dataset identifier in the second queue, it means that the text training dataset corresponding to the to-be-obtained training dataset identifier is accessed by the server for the first time. At this time, the server obtains the storage path of the text training dataset corresponding to the to-be-obtained training dataset identifier according to the to-be-obtained training dataset identifier, and obtains the download address of the to-be-obtained training dataset according to this storage path.
[0133] Step 504, download the target text training dataset based on the download address of the to-be-obtained training dataset, store the target text training dataset on the disk, and obtain the disk storage address and the current access count.
[0134] The disk storage address refers to the specific storage location of the target text training dataset on the disk. The current access count refers to the access count of the target text training dataset, that is, the number of times the target text training dataset is read by the server. At this time, the current access count is the initial value.
[0135] Specifically, the server uses the download address of the to-be-obtained training dataset to download data, obtains the target text training dataset, then stores the target text training dataset on the disk of the server, and obtains the disk storage address and the current access count.
[0136] Step 506, write the to-be-obtained training dataset identifier, the current access count, and the disk storage address into the second queue.
[0137] Specifically, the server writes the to-be-obtained training dataset identifier, the current access count, and the disk storage address into the second queue in an associated manner.
[0138] In one embodiment, as Figure 6 shown, step 506, that is, writing the to-be-obtained training dataset identifier, the current access count, and the disk storage address into the second queue, includes:
[0139] Step 602, when the second queue is full, find the to-be-deleted second text training dataset identifier corresponding to the target time point in the second queue. The target time point refers to the oldest time point.
[0140] Among them, the second queue is set with an upper limit for storing data. For example, it stores the relevant information of 100 data sets. The target time point refers to the oldest time point in the second queue, that is, the last time point in the second queue, which is the historical time point with the longest distance from the current time point. The second text training data set identifier to be deleted refers to the second text training data set identifier that needs to be deleted.
[0141] Specifically, when the second queue reaches the storage upper limit, it means the second queue is full. When the relevant information of a new text training data set needs to be written, this relevant information refers to the training data set identifier to be obtained, the current access times, and the disk storage address. At this time, the second text training data set identifier to be deleted corresponding to the target time point is searched for in the second queue.
[0142] Step 604: Delete the second text training data set identifier to be deleted corresponding to the target time point, the historical access times to be deleted corresponding to the second text training data set identifier to be deleted, and the disk storage address to be deleted from the second queue.
[0143] Among them, the second text training data set identifier to be deleted refers to the second text training data set identifier corresponding to the target time point, which is the second text training data set identifier that needs to be deleted. The historical access times to be deleted refer to the historical access times corresponding to the target time point, which are the historical access times that need to be deleted. The disk storage address to be deleted refers to the disk storage address corresponding to the target time point, which is the disk storage address that needs to be deleted.
[0144] Specifically, the server deletes the data associated with the target time point in the second queue. That is, it deletes the second text training data set identifier to be deleted corresponding to the target time point, the historical access times to be deleted corresponding to the second text training data set identifier to be deleted, and the disk storage address to be deleted from the second queue.
[0145] Step 606: Write the training data set identifier to be obtained, the current access times, and the disk storage address into the second queue.
[0146] Specifically, since the data associated with the target time point has been deleted from the second queue, at this time, the server can write the training data set identifier to be obtained, the current access times, and the disk storage address into the second queue.
[0147] In a specific embodiment, such as Figure 7As shown in the figure, it is a schematic diagram of the updated process of the second queue. Specifically: The sizes of the first queue and the second queue are preset in the server. Then, when data access is performed, first query whether the dataset to be accessed is in the first queue. When it is in the first queue, read the dataset from the first queue and update the data access time. When it is not in the first queue, update the second queue through the LRU (Least Recently Used) policy, that is, search for the dataset in the second queue and update the access count of the dataset in the second queue. When the access count exceeds the threshold, delete the relevant information of the dataset from the second queue, write the dataset into the first queue, and update the first queue using the LRU policy, then read the dataset and return. When the access count does not exceed the threshold, read the dataset from the disk and return. As Figure 8 As shown in the figure, it is a schematic diagram of the interaction between queues. Among them, when it is found in the first queue, the first queue is updated through the LRU policy. When it is not found in the first queue, the second queue is updated through the LRU policy. When the access count of the dataset in the second queue is greater than the threshold, it will be written into the first queue, and the first queue is updated using the LRU policy.
[0148] In the above embodiment, by processing the second queue, relevant information of the text training dataset that has been accessed can be saved in the second queue. When subsequent access is performed, relevant information can be directly found in the second queue, so that the text training dataset to be obtained can be quickly read, improving the efficiency.
[0149] In one embodiment, as Figure 9 shown in the figure, step 210, writing into the first queue includes:
[0150] Step 902, when the first queue is full, search for the identifier of the first text training dataset to be deleted corresponding to the target time point in the first queue. The target time point refers to the time point with the oldest time.
[0151] Among them, the first queue is set with an upper limit for storing data. For example, it stores 100 datasets, each dataset stores the first 100 files, and each file stores the first 100 lines of text data. The target time point refers to the time point with the oldest time in the first queue, that is, the time point at the end of the first queue, that is, the historical time point with the longest distance from the current time point. The identifier of the first text training dataset to be deleted refers to the identifier of the first text training dataset that needs to be deleted.
[0152] Specifically, when the first queue reaches its storage limit, it indicates that the first queue is full. When it is necessary to write the relevant information of the new text training dataset, this relevant information refers to the current access time point, the target text training dataset identifier, and the target text training dataset corresponding to the target text training dataset identifier. At this time, search for the identifier of the first text training dataset to be deleted corresponding to the target time point in the first queue.
[0153] Step 904: Delete the identifier of the first text training dataset to be deleted corresponding to the target time point and the first text training dataset to be deleted corresponding to the identifier of the first text training dataset to be deleted from the first queue.
[0154] Among them, the identifier of the first text training dataset to be deleted refers to the identifier of the first text training dataset corresponding to the target time point, which is the identifier of the first text training dataset that needs to be deleted. The first text training dataset to be deleted refers to the first text training dataset corresponding to the target time point, which is the first text training dataset that needs to be deleted.
[0155] Specifically, the server deletes the data associated with the target time point in the first queue. That is, it deletes the identifier of the first text training dataset to be deleted corresponding to the target time point and the first text training dataset to be deleted corresponding to the identifier of the first text training dataset to be deleted from the first queue.
[0156] Step 906: Obtain the current access time point, and write the current access time point, the target text training dataset identifier, and the target text training dataset corresponding to the target text training dataset identifier into the first queue.
[0157] Specifically, the server obtains the current access time point, the target text training dataset identifier, and the target text training dataset identifier to be written into the first queue. The access times of the target text training datasets corresponding to the target text training dataset identifier meet the caching conditions. Then the server writes the current access time point, the target text training dataset identifier, and the target text training dataset corresponding to the target text training dataset identifier into the first queue in an associated manner.
[0158] In a specific embodiment, as Figure 10 shown, it is a schematic flow diagram for the use of the first queue. Among them, the server pre-sets the size of the first queue, and then determines whether the dataset is in the first queue. When it is in the first queue, read the dataset from the first queue and update the data access time. When it is not in the first queue, first determine whether the first queue is full, that is, whether it can continue to write the dataset. When the first queue is full, use the LRU strategy to delete the data in the first queue, that is, delete the least recently used time. When the first queue is not full, insert the dataset into the first queue and update the data access time. As Figure 11As shown, it is a schematic diagram of data processing for the first queue. Among them, a linked list is used to represent the priority of data, and the head of the linked list (the bottom in the figure) has the lowest priority. The data is stored in dictionary form. Initially, the first queue is empty, and then the data sets 17, 0, and 11 are written in sequence. At this time, the first queue is full. When it is necessary to continue writing the data set 12, the data set 17 with the lowest priority is deleted, and then the data set 12 is inserted. When it is necessary to continue accessing the data set 0, the priority of the data set 0 in the first queue is updated to the highest. Among them, the priority can be determined according to the access time. The newer the time, the higher the priority.
[0159] In the above embodiment, by processing the first column, relevant information of the text training data set that meets the cache condition can be saved in the first column. When accessing later, the relevant information of the text training data set that meets the cache condition can be directly found from the first queue, so that the text training data set to be obtained can be quickly read, improving the efficiency.
[0160] In one embodiment, after step 202, that is, after receiving the training data acquisition request, and the training data acquisition request carries the identifier of the training data set to be obtained, the following steps are further included:
[0161] When the type of the training data set to be obtained corresponding to the identifier of the training data set to be obtained is a picture type, the corresponding picture download address is obtained according to the picture training data set identifier, and the picture training data set is downloaded based on the picture download address to obtain the picture training data set; the picture training data set is input into the network model for training.
[0162] Among them, the picture download address refers to the URL (Uniform Resource Locator) of the picture. The picture training data set refers to the set of pictures used during training.
[0163] Specifically, the server parses the data acquisition request to obtain the data format to be obtained. When the data format is a picture format, it indicates that the type of the training data set to be obtained corresponding to the identifier of the training data set to be obtained is a picture type. At this time, the server obtains the storage path of the picture according to the picture training data set identifier, constructs the full path, obtains the picture download address according to the full path, and then uses the picture download address to download to obtain the picture training data set. Then the picture training data set is input into the network model for training to obtain a trained picture network model, and this picture network model is used to process the input picture according to the model task type. That is, by obtaining the picture download address to download and obtain the picture training data set, and then using the picture training data to train the network model, the efficiency of training the network model can be improved.
[0164] In one embodiment, as Figure 12As shown, a data display method is provided. Taking the terminal in Figure 1 as an example for illustration, it can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. It includes the following steps:
[0165] Step 1202: Receive a data display instruction sent through a data display page, and send a data acquisition request to the server based on the data display instruction. The data acquisition request carries an identifier of a dataset to be displayed. The server is used to, when the type of the dataset to be displayed corresponding to the identifier of the dataset to be displayed is a text type, search for a corresponding target text dataset identifier in the first queue. The first queue stores various first text dataset identifiers and corresponding first text datasets. When the corresponding target text dataset identifier is not found in the first queue, search for the target text dataset identifier corresponding to the identifier of the dataset to be displayed in the second queue. The second queue stores various second text dataset identifiers, corresponding historical access times, and corresponding disk storage locations. Both the first queue and the second queue are located in the data cache. When the corresponding target text dataset identifier is found in the second queue, obtain the target disk storage location corresponding to the found target text dataset identifier, read the target text dataset corresponding to the target text dataset identifier from the corresponding disk based on the target disk storage location, update the historical access times corresponding to the target text dataset identifier in the second queue. The second text datasets corresponding to the second text dataset identifiers in the second queue are used to be written into the first queue when the corresponding historical access times meet the cache condition, and return the target text dataset to the terminal.
[0166] Among them, the data display page refers to the page in the terminal for data display. The type of the dataset to be displayed refers to the type of data to be displayed, including text type and picture type. The text type refers to the data displayed in text form, and the picture type refers to the data displayed in picture form. The target text dataset identifier refers to the identifier of the text dataset to be displayed. The identifier of the dataset to be displayed uniquely identifies the dataset that needs to be displayed. The target disk storage location refers to the specific storage location of the target text dataset on the disk. The first text dataset identifier is used to identify the dataset in the first queue, and the first text dataset refers to the text dataset in the first queue. The second text dataset identifier is used to identify the dataset in the second queue. The historical access times refer to the number of times the second text dataset corresponding to the second text dataset identifier has been accessed historically. The disk storage location refers to the storage location of the second text dataset corresponding to the second text dataset identifier in the second queue. The target text dataset refers to the text dataset to be displayed.
[0167] Specifically, the terminal receives a data display instruction sent through the data display page, and then sends a data acquisition request to the server based on the data display instruction. The data acquisition request carries an identifier of the data set to be displayed. The server receives the data acquisition request, parses the data acquisition request, and obtains the identifier of the data set to be displayed. When the type of the data set to be displayed corresponding to the identifier of the data set to be displayed is a text type, the server searches for the corresponding target text data set identifier in the first queue based on the identifier of the data set to be displayed. When the corresponding target text data set identifier is not found in the first queue, the server searches for the target text data set identifier corresponding to the identifier of the data set to be displayed in the second queue. Both the first queue and the second queue are located in the data cache. When the corresponding target text data set identifier is found in the second queue, the server obtains the target disk storage location corresponding to the found target text data set identifier, reads the target text data set corresponding to the target text data set identifier from the corresponding disk based on the target disk storage location, updates the historical access times corresponding to the target text data set identifier in the second queue. The second text data set corresponding to the second text data set identifier in the second queue is used to be written into the first queue when the corresponding historical access times meet the cache condition. Finally, the server returns the target text data set to the terminal.
[0168] Step 1204, obtain the target text data set returned by the server, and display the target text data set through the data display page.
[0169] Specifically, the terminal obtains the target text data set returned by the server, and displays the target text data set through the data display page.
[0170] The above data display method, device, computer device, and storage medium. The terminal receives a data display instruction sent through a data display page, and based on the data display instruction, sends a data acquisition request to the server. The data acquisition request carries an identifier of a dataset to be displayed. Then, when the type of the dataset to be displayed corresponding to the identifier of the dataset to be displayed is a text type, the server is used to search for a corresponding target text dataset identifier in a first queue based on the identifier of the dataset to be displayed. When the corresponding target text dataset identifier is not found in the first queue, it searches for the target text dataset identifier corresponding to the identifier of the dataset to be displayed in a second queue. Both the first queue and the second queue are located in the data cache. When the corresponding target text dataset identifier is found in the second queue, it obtains the target disk storage location corresponding to the found target text dataset identifier, reads the target text dataset corresponding to the target text dataset identifier from the corresponding disk based on the target disk storage location, and returns the target text dataset to the terminal; obtains the target text dataset returned by the server, and displays the target text dataset through the data display page. That is, through the first queue and the second queue, multi-level caching is used for access when obtaining display data, making full use of the cache and the disk to reduce the time consumed by downloading, parsing, and transmission, etc., avoiding the server from downloading the data to be displayed from the storage medium, so that the data to be displayed can be quickly obtained and the data display speed can be improved.
[0171] In one embodiment, as Figure 13 shown, the method further includes:
[0172] Step 1302, receive a data display instruction sent through a data display page, and based on the data display instruction, send a data synchronization request to the server. The data synchronization request carries an identifier of a dataset to be displayed. The server responds to the data synchronization request, generates an asynchronous task identifier, and records the asynchronous task state corresponding to the asynchronous task identifier as unfinished, and returns the asynchronous task identifier to the terminal; the server is used to execute the asynchronous task corresponding to the asynchronous task identifier. The asynchronous task refers to the task of the server to obtain a target text dataset when the type of the dataset to be displayed corresponding to the identifier of the dataset to be displayed is a text type. Obtaining the target text dataset means searching for the corresponding text dataset identifier in the first queue using the identifier of the dataset to be displayed. When the corresponding text dataset identifier is not found in the first queue, searching for the text dataset identifier corresponding to the identifier of the dataset to be displayed in the second queue. When the corresponding text dataset identifier is found in the second queue, obtaining the disk storage location corresponding to the found text dataset identifier, reading the text dataset corresponding to the text dataset identifier from the corresponding disk based on the disk storage location, updating the historical access times corresponding to the text dataset identifier in the second queue, and updating the asynchronous task state corresponding to the asynchronous task identifier to completed.
[0173] Among them, the asynchronous task identifier is used to uniquely identify an asynchronous task, which refers to the task of obtaining the data set to be displayed, that is, the task of reading the target text data set needs to be completed through asynchronous task scheduling. The asynchronous task status refers to the status of the execution of the asynchronous task, including two statuses: unfinished and completed.
[0174] Specifically, the terminal receives a data display instruction sent through the data display page, and the terminal sends a data synchronization request to the server based on the data display instruction. The data synchronization request carries an identifier of the data set to be displayed. When the server receives the data synchronization request, it responds to the data synchronization request, generates an asynchronous task identifier, records the asynchronous task status corresponding to the asynchronous task identifier as unfinished, and returns the asynchronous task identifier to the terminal.
[0175] Then the server executes the asynchronous task corresponding to the asynchronous task identifier, that is, the server searches for the corresponding text data set identifier in the first queue by using the identifier of the data set to be displayed parsed from the data synchronization request. When the corresponding text data set identifier is not found in the first queue, it searches for the text data set identifier corresponding to the identifier of the data set to be displayed in the second queue. When the corresponding text data set identifier is found in the second queue, it obtains the disk storage location corresponding to the found text data set identifier, reads the text data set corresponding to the text data set identifier from the corresponding disk based on the disk storage location, updates the historical access times corresponding to the text data set identifier in the second queue, and updates the asynchronous task status corresponding to the asynchronous task identifier to completed.
[0176] Step 1304: Obtain the asynchronous task identifier returned by the server, and send an asynchronous request to the server at a preset time interval based on the asynchronous task identifier. The server responds to the asynchronous request, queries the asynchronous task status, and when the asynchronous task status is completed, returns the target text data set corresponding to the identifier of the data set to be displayed to the terminal.
[0177] Among them, the preset time interval refers to the time interval set in advance.
[0178] Specifically, when the terminal obtains the asynchronous task identifier returned by the server, it continuously polls and sends an asynchronous request to the server using the asynchronous task identifier, that is, sends an asynchronous request to the server at a preset time interval. When the server receives the asynchronous request, it responds to the asynchronous request and queries the asynchronous task status. When the asynchronous task status is unfinished, it returns the information that the asynchronous task status is unfinished to the terminal. When the asynchronous task status is completed, it obtains the target text data set corresponding to the identifier of the data set to be displayed read when completed, and then returns the target text data set corresponding to the identifier of the data set to be displayed to the terminal.
[0179] Step 1306: Obtain the target text data set returned by the server, and display the target text data set through the data display page.
[0180] Specifically, the terminal obtains the target text dataset returned by the server and displays the target text dataset through the data display page.
[0181] In the above embodiment, the target text dataset is obtained and displayed through an asynchronous request, so that resources such as occupied threads can be released to avoid blocking. And the results of multiple calls can be uniformly returned to the terminal, that is, the read target text dataset is returned to the terminal for display, which is convenient and fast.
[0182] In one embodiment, the data acquisition request carries an identifier of the dataset to be displayed and an identifier of the requestor.
[0183] Step 1202, sending a data acquisition request to the server based on the data display instruction, including the steps of:
[0184] Sending a target data acquisition request to the server based on the data display instruction, the server parses the target data acquisition request to obtain the identifier of the requestor, and matches the identifier of the requestor with a preset data reading permission list. When there is a matching identifier of the requestor, the server responds to the data synchronization request, generates an asynchronous task identifier, records the asynchronous task status corresponding to the asynchronous task identifier as unfinished, and returns the asynchronous task identifier to the terminal.
[0185] Among them, the identifier of the requestor is used to uniquely identify the terminal capable of reading data. The preset data reading permission list refers to a list of terminals with data reading permissions set in advance, and each authorized identifier of the requestor is stored in the preset data reading permission list.
[0186] Specifically, the terminal sends a target data acquisition request to the server based on the data display instruction. The server receives the target data acquisition request, parses the target data acquisition request to obtain the identifier of the requestor, and then matches the identifier of the requestor with the preset data reading permission list. When there is a matching identifier of the requestor, it means that the terminal has the permission to read data. At this time, the server responds to the data synchronization request, generates an asynchronous task identifier, records the asynchronous task status corresponding to the asynchronous task identifier as unfinished, and returns the asynchronous task identifier to the terminal. That is, the permission to read data is restricted through the preset data reading permission list, thereby improving the security of data reading.
[0187] In one embodiment, as Figure 14 shown, the method further includes:
[0188] Step 1402: Receive a data display instruction sent through the data display page, and send a data acquisition request to the server based on the data display instruction. The data acquisition request carries an identifier of the dataset to be displayed. The server is configured to, when the type of the dataset to be displayed corresponding to the identifier of the dataset to be displayed is a text type, search for the corresponding target text dataset identifier in the first queue based on the identifier of the dataset to be displayed. When the corresponding target text dataset identifier is found in the first queue, obtain the target cache storage location corresponding to the target text dataset identifier, read the target text dataset corresponding to the target text dataset identifier from the corresponding cache based on the target cache storage location, obtain the current time point, update the first access time point corresponding to the target text dataset identifier based on the current time point, and return the target text dataset corresponding to the identifier of the dataset to be displayed to the terminal.
[0189] Wherein, the current time point refers to the time point when accessing data currently, that is, the time point when reading data. The first access time point refers to the access time point corresponding to the target text dataset identifier in the first queue.
[0190] Specifically, the terminal receives a data display instruction sent through the data display page, and sends a data acquisition request to the server based on the data display instruction. The data acquisition request carries an identifier of the dataset to be displayed. The server receives the data acquisition request, parses the data acquisition request to obtain the identifier of the dataset to be displayed, and then, when the type of the dataset to be displayed corresponding to the identifier of the dataset to be displayed is a text type, searches for the corresponding target text dataset identifier in the first queue based on the identifier of the dataset to be displayed. When the corresponding target text dataset identifier is found in the first queue, obtain the target cache storage location corresponding to the target text dataset identifier, read the target text dataset corresponding to the target text dataset identifier from the corresponding cache based on the target cache storage location. The server obtains the current time point, updates the first access time point corresponding to the target text dataset identifier based on the current time point, and returns the target text dataset corresponding to the identifier of the dataset to be displayed to the terminal.
[0191] Step 1404: Obtain the target text dataset returned by the server, and display the target text dataset through the data display page.
[0192] Specifically, the terminal obtains the target text dataset returned by the server, and displays the target text dataset through the data display page.
[0193] In the above embodiment, by directly reading the target text dataset from the cache storage location in the first queue, the reading within constant time is realized, and then the read target text dataset is returned to the terminal, and the terminal displays it through the data display page, improving the display efficiency of the text data.
[0194] In a specific embodiment, such asFigure 15 As shown in the figure, it is a schematic diagram of a specific framework for text dataset display. Among them, the front end sends a data display synchronization request to the server. The data display synchronization request carries a text dataset id (Identity document, unique encoding). The server authenticates through the gateway, that is, matches the requester identity through a pre-set permission list. When the match is consistent, that is, the gateway authentication passes. At this time, the server sends the data display synchronization request to the data center server through routing. The data center server responds to the synchronization request to generate an asynchronous task id and returns the asynchronous task id to the terminal. The server performs asynchronous task scheduling and triggers the task execution logic, that is, obtains the target text dataset through multi-level cache access using the first queue and the second queue. When the task execution is completed, the task status corresponding to the asynchronous task id in the asynchronous task status table is updated to completed, that is, the status is updated from not completed to completed. When the front end receives the asynchronous task id, it continuously polls and sends a dataset display asynchronous request, and always obtains the returned status 0 during the execution of the server asynchronous task, that is, the asynchronous task is not completed. When the asynchronous status is updated to completed, the terminal obtains the target text dataset and status 1 in the form of a binary stream after file parsing, that is, the asynchronous task is completed. Then the text content in the file is displayed on the front end.
[0195] In one embodiment, as Figure 16 shown, the method further includes:
[0196] Step 1602: Receive a data display instruction sent through the data display page, and send a data acquisition request to the server based on the data display instruction. The data acquisition request carries an identifier of the dataset to be displayed. When the type of the dataset to be displayed corresponding to the identifier of the dataset to be displayed is a picture type, the server obtains the corresponding picture storage address according to the picture dataset identifier, converts the picture storage address to obtain a picture download address, and returns the picture download address to the terminal.
[0197] Among them, the picture storage address refers to the specific storage location of the picture data. The picture download address is used to download the picture data.
[0198] Specifically, when it is necessary to display a picture dataset, the terminal receives a data display instruction sent through the data display page, and sends a data acquisition request to the server based on the data display instruction. The data acquisition request carries an identifier of the dataset to be displayed. The server parses the received data acquisition request to obtain the identifier of the dataset to be displayed. Then, when the format of the dataset to be displayed corresponding to the identifier of the dataset to be displayed is a picture format, it is obtained that the type of the dataset to be displayed corresponding to the identifier of the dataset to be displayed is a picture type. At this time, the server obtains the corresponding picture storage address according to the picture dataset identifier, converts the picture storage address to obtain a picture download address, and returns the picture download address to the terminal.
[0199] Step 1604: Receive the image download address returned by the server, and download the image based on the image download address to obtain an image data set.
[0200] Step 1606: Display the image data set on the data display page.
[0201] Specifically, the terminal receives the image download address returned by the server, downloads the image based on the image download address to obtain an image data set, and displays the image data set on the data display page.
[0202] In the above embodiment, the terminal obtains the image download address through the server, then downloads and displays the image, avoiding the time consumption of the server downloading the image and then transmitting it, and improving the image display efficiency.
[0203] In one embodiment, for the data display method disclosed in the present application, the text data set and the image data set can be stored on the blockchain.
[0204] In a specific embodiment, as Figure 17 shown, it is a schematic diagram of the specific framework for displaying the image data set. Among them, the front end sends a data display synchronization request to the server. The data display synchronization request carries the image data set id. The server performs gateway authentication, that is, matches the requester identifier through a pre-set permission list. When the match is consistent, that is, the gateway authentication passes. At this time, the server sends the data display synchronization request to the data center server through routing. The data center server responds to the synchronization request to generate an asynchronous task id and returns the asynchronous task id to the terminal. The server performs asynchronous task scheduling and triggers the task execution logic. According to the image data set id, it obtains the data set file storage path, then constructs the full path, generates the image download address according to the full path, returns the image download address to the terminal, and updates the task status corresponding to the asynchronous task id in the asynchronous task status table to completed, that is, updates the status from not completed to completed. When the front end receives the asynchronous task id, it continuously polls and sends a data set display asynchronous request, and always obtains the returned status 0 during the execution of the server asynchronous task, that is, the asynchronous task is not completed. When the asynchronous status is updated to completed, the terminal obtains the image download address and status 1, that is, the asynchronous task is completed. Then the front end downloads the image according to the image download address and displays the downloaded image on the front end. That is, by returning the image download address to the terminal, the terminal is enabled to download, thus avoiding the time consumption of the server downloading the image and transmitting it, and improving the image display efficiency.
[0205] In a specific embodiment, as Figure 18 shown, it is a method for processing network model data, specifically including:
[0206] Step 1802: Receive a training data acquisition request, which carries the identifier of the training data set to be acquired.
[0207] Step 1804: When the type of the training data set to be acquired corresponding to the identifier of the training data set to be acquired is of text type, search for the corresponding target text training data set identifier in the first queue based on the identifier of the training data set to be acquired. When the corresponding target text training data set identifier is not found in the first queue, search for the target text training data set identifier corresponding to the identifier of the training data set to be acquired in the second queue.
[0208] Step 1806: When the corresponding target text training data set identifier is found in the second queue, obtain the corresponding target disk storage location of the found target text training data set identifier. Read the target text training data set corresponding to the target text training data set identifier from the corresponding disk based on the target disk storage location. Input the target text training data set into the network model for training.
[0209] Step 1808: When the type of the training data set to be acquired corresponding to the identifier of the training data set to be acquired is of picture type, obtain the corresponding picture download address according to the picture training data set identifier, and perform a download based on the picture download address to obtain the picture training data set. Input the picture training data set into the network model for training.
[0210] This application also provides an application scenario, which applies the above data display method. Specifically,
[0211] The data display method of this application is applied in the data center system of the intelligent learning platform. As Figure 19 shown, it is a framework schematic diagram of the data center system, which includes two major parts: dataset management and storage management. Storage management mainly completes storage adaptation, liveness detection, expiration cleaning, etc.; dataset management mainly completes write-type operations such as import and release, and detail-type operations such as download and details. After the artificial intelligence algorithm-related platforms such as the annotation platform, training platform, and inference platform transfer the dataset to the specified medium through import, they perform artificial intelligence calculations such as annotation, training, and inference. For example, the text dataset can be processed through natural language processing technology, and the picture dataset can be processed through computer vision technology. Users can view the training data of the trained artificial intelligence model through the intelligent learning platform. That is, the data can be displayed through the data details page. As Figure 20As shown in the figure, it is a schematic diagram of the specific process for data display in the data center system. Among them, the sizes of the first queue and the second queue in the cache are set in the data center system. When it is necessary to display the text data set, the data set ID to be displayed is obtained. It is judged whether it is the first access according to the data set ID. When it is not the first access, when the data set ID is found in the first queue, the data volume size of the data set to be displayed is obtained. When the data volume size of the data set to be displayed exceeds the data volume cached in the first queue, the file is read from the disk and parsed into binary. Then, the parsed binary stream is returned to the terminal. When the data volume size of the displayed data set does not exceed the data volume cached in the first queue, the cached binary stream is directly read from the cache storage location, and the parsed binary stream is returned to the terminal. When it is the first access, the file download address of the text data set is obtained, downloaded using the data set file download address, and stored on the disk. The file is parsed into a binary stream, and the second queue and the first queue are established. Among them, both the first queue and the second queue use the LRU policy as the update policy. Then, the parsed binary stream is returned to the terminal. As Figure 21 shown, it is a schematic diagram of part of the data of the text data set to be displayed. Among them, when it is necessary to display the picture data set, the picture download address returned by the server is obtained, and then the picture is downloaded using the picture download address. As Figure 22 shown, it is a schematic diagram of part of the data of the picture data set.
[0212] It should be understood that although Figure 2-20 the steps in the flowchart in [document] are displayed in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 2-20 at least a part of the steps in the flowchart in [document] may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of the steps or stages in other steps or other steps.
[0213] In one embodiment, as Figure 23 shown, a network model data processing device 2300 is provided. This device can be a software module, a hardware module, or a combination of both to become a part of a computer device. Specifically, this device includes: a request receiving module 2302, a first search module 2304, a second search module 2306, a location acquisition module 2308, a data reading module 2310, and a training module 2312, where:
[0214] A request receiving module 2302, configured to receive a training data acquisition request, where the training data acquisition request carries an identifier of a training data set to be acquired;
[0215] A first search module 2304, configured to, when the type of the training data set to be acquired corresponding to the identifier of the training data set to be acquired is a text type, search for a corresponding target text training data set identifier in a first queue; the first queue stores various first text training data set identifiers and corresponding first text training data sets;
[0216] A second search module 2306, configured to, when the corresponding target text training data set identifier is not found in the first queue, search for the target text training data set identifier corresponding to the identifier of the training data set to be acquired in a second queue; the second queue stores various second text training data set identifiers, corresponding historical access times, and corresponding disk storage locations, and both the first queue and the second queue are located in a data cache;
[0217] A location acquisition module 2308, configured to, when the corresponding target text training data set identifier is found in the second queue, acquire the target disk storage location corresponding to the found target text training data set identifier;
[0218] A data reading module 2310, configured to read the target text training data set corresponding to the target text training data set identifier from a corresponding disk based on the target disk storage location, update the historical access time corresponding to the target text training data set identifier in the second queue, and the second text training data set corresponding to the second text training data set identifier in the second queue is used to be written into the first queue when the corresponding historical access time meets a cache condition;
[0219] A training module 2312, configured to input the target text training data set into a network model for training, where the network model is used to process input data according to a model task type.
[0220] In one embodiment, the first queue stores various first text training data set identifiers, corresponding first text training data sets, corresponding cache storage locations, and corresponding first access time points;
[0221] The network model data processing device 2300 further includes: a first acquisition module, configured to, when the corresponding target text training data set identifier is found in the first queue, acquire the target cache storage location corresponding to the target text training data set identifier, read the target text training data set corresponding to the target text training data set identifier from a corresponding cache based on the target cache storage location; acquire the current time point, and update the first access time point corresponding to the target text training data set identifier based on the current time point.
[0222] In one embodiment, the training data acquisition request further carries the data volume of the training data set to be acquired; the first acquisition module is further configured to, when the corresponding target text training data set identifier is found in the first queue, acquire the cached data volume of the target text training data set corresponding to the target text training data set identifier in the first queue; when the cached data volume exceeds the data volume of the training data set to be acquired, obtain the target cache storage location corresponding to the target text training data set identifier, and read the target text training data set with the data volume of the training data set to be acquired from the corresponding cache based on the target cache storage location.
[0223] In one embodiment, the first acquisition module is further configured to, when the cached data volume does not exceed the data volume of the training data set to be acquired, obtain the corresponding target disk storage location according to the target text training data set identifier, and read the target text training data set with the data volume of the training data set to be acquired from the corresponding disk based on the target disk storage location.
[0224] In one embodiment, the network model data processing device 2300 further includes:
[0225] The download module is configured to, when the corresponding target text training data set identifier of the training data set to be acquired is not found in the second queue, acquire the download address of the training data set to be acquired based on the identifier of the training data set to be acquired; download the target text training data set based on the download address of the training data set to be acquired, store the target text training data set in the disk, and obtain the disk storage address and the current access times; write the identifier of the training data set to be acquired, the current access times, and the disk storage address into the second queue.
[0226] In one embodiment, the download module is further configured to, when the second queue is full, find the identifier of the second text training data set to be deleted corresponding to the target time point in the second queue, where the target time point refers to the oldest time point; delete the identifier of the second text training data set to be deleted corresponding to the target time point, the historical access times to be deleted corresponding to the identifier of the second text training data set to be deleted, and the disk storage address to be deleted from the second queue; write the identifier of the training data set to be acquired, the current access times, and the disk storage address into the second queue.
[0227] In one embodiment, the data reading module 2310 is further configured to, when the first queue is full, find the identifier of the first text training data set to be deleted corresponding to the target time point in the first queue, where the target time point refers to the oldest time point; delete the identifier of the first text training data set to be deleted corresponding to the target time point and the first text training data set corresponding to the identifier of the first text training data set to be deleted from the first queue; obtain the current access time point, and write the current access time point, the target text training data set identifier, and the target text training data set corresponding to the target text training data set identifier into the first queue.
[0228] In one embodiment, the network model data processing device 2300 further includes:
[0229] An image download module, configured to, when the type of the to-be-acquired training data set corresponding to the to-be-acquired training data set identifier is an image type, obtain a corresponding image download address according to the image training data set identifier, perform a download based on the image download address to obtain an image training data set; and input the image training data set into the network model for training.
[0230] In one embodiment, as Figure 24 shown, a data display device 2400 is provided. The device can be a software module, a hardware module, or a combination of both to form a part of a computer device. The device specifically includes: a request sending module 2402 and a text display module 2404, where:
[0231] The request sending module 2402 is configured to receive a data display instruction sent through a data display page, and send a data acquisition request to a server based on the data display instruction. The data acquisition request carries a to-be-displayed data set identifier. The server is configured to, when the type of the to-be-displayed data set corresponding to the to-be-displayed data set identifier is a text type, search for a corresponding target text data set identifier in a first queue. The first queue stores each first text data set identifier and the corresponding first text data set. When the corresponding target text data set identifier is not found in the first queue, search for the target text data set identifier corresponding to the to-be-displayed data set identifier in a second queue. The second queue stores each second text data set identifier, the corresponding historical access times, and the corresponding disk storage location. Both the first queue and the second queue are located in a data cache. When the corresponding target text data set identifier is found in the second queue, obtain the target disk storage location corresponding to the found target text data set identifier, read the target text data set corresponding to the target text data set identifier from the corresponding disk based on the target disk storage location, update the historical access times corresponding to the target text data set identifier in the second queue. The second text data set corresponding to the second text data set identifier in the second queue is configured to be written into the first queue when the corresponding historical access times meet the cache condition, and return the target text data set to the terminal;
[0232] The text display module 2404 is configured to obtain the target text data set returned by the server and display the target text data set through the data display page.
[0233] In one embodiment, the data display device 2400 further includes:
[0234] An asynchronous execution module is used to receive a data display instruction sent through a data display page, send a data synchronization request to a server based on the data display instruction, where the data synchronization request carries an identifier of a data set to be displayed. The server responds to the data synchronization request, generates an asynchronous task identifier, records the asynchronous task status corresponding to the asynchronous task identifier as unfinished, and returns the asynchronous task identifier to the terminal. The server is used to execute the asynchronous task corresponding to the asynchronous task identifier. The asynchronous task refers to a task where when the data set type to be displayed corresponding to the identifier of the data set to be displayed is a text type, the server obtains a target text data set. Obtaining the target text data set means searching for the corresponding text data set identifier in the first queue using the identifier of the data set to be displayed. When the corresponding text data set identifier is not found in the first queue, searching for the text data set identifier corresponding to the identifier of the data set to be displayed in the second queue. When the corresponding text data set identifier is found in the second queue, obtaining the disk storage location corresponding to the found text data set identifier, reading the text data set corresponding to the text data set identifier from the corresponding disk based on the disk storage location, updating the historical access times corresponding to the text data set identifier in the second queue, and updating the asynchronous task status corresponding to the asynchronous task identifier as completed. Obtaining the asynchronous task identifier returned by the server, and sending an asynchronous request to the server at a preset time interval based on the asynchronous task identifier. The server responds to the asynchronous request, queries the asynchronous task status, and when the asynchronous task status is completed, returns the target text data set corresponding to the identifier of the data set to be displayed to the terminal. Obtaining the target text data set returned by the server, and displaying the target text data set through the data display page.
[0235] In one embodiment, the data acquisition request carries an identifier of a data set to be displayed and an identifier of the requestor.
[0236] The request sending module 2402 is further configured to send a target data acquisition request to the server based on the data display instruction. The server parses the target data acquisition request to obtain the identifier of the requestor, and matches the identifier of the requestor with a preset data reading permission list. When there is a matching identifier of the requestor, the server responds to the data synchronization request, generates an asynchronous task identifier, records the asynchronous task status corresponding to the asynchronous task identifier as unfinished, and returns the asynchronous task identifier to the terminal.
[0237] In one embodiment, the data display device 2400 further includes:
[0238] A time update module, configured to receive a data display instruction sent through a data display page, and send a data acquisition request to a server based on the data display instruction, where the data acquisition request carries an identifier of a data set to be displayed; the server is configured to, when the type of the data set to be displayed corresponding to the identifier of the data set to be displayed is a text type, search for a corresponding target text data set identifier in a first queue based on the identifier of the data set to be displayed, and when the corresponding target text data set identifier is found in the first queue, obtain a target cache storage location corresponding to the target text data set identifier, read a target text data set corresponding to the target text data set identifier from the corresponding cache based on the target cache storage location; obtain a current time point, update a first access time point corresponding to the target text data set identifier based on the current time point, and return the target text data set corresponding to the identifier of the data set to be displayed to a terminal; obtain the target text data set returned by the server, and display the target text data set through the data display page.
[0239] In one embodiment, the data display device 2400 further includes:
[0240] A picture display module, configured to receive a data display instruction sent through a data display page, and send a data acquisition request to a server based on the data display instruction, where the data acquisition request carries an identifier of a data set to be displayed; when the type of the data set to be displayed corresponding to the identifier of the data set to be displayed is a picture type, the server obtains a corresponding picture storage address according to the picture data set identifier, converts the picture storage address to obtain a picture download address, and returns the picture download address to the terminal; receive the picture download address returned by the server, download a picture based on the picture download address to obtain a picture data set; and display the picture data set on the data display page.
[0241] For the specific definitions of the network model data processing device and the data display device, reference may be made to the definitions of the network model data processing and data display methods in the foregoing text, which will not be elaborated here. Each module in the above-mentioned network model data processing device and data display device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.
[0242] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 25As shown in the figure. The computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store text and picture data sets. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements a network model data processing method and a data display method.
[0243] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 26 shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a network model data processing method and a data display method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0244] Those skilled in the art can understand that Figure 25 and Figure 26 the structures shown in the figure are only block diagrams of some structures related to the solution of this application, and do not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0245] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.
[0246] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which when executed by a processor implements the steps in the above method embodiments.
[0247] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.
[0248] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium, and when the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0249] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0250] The above-described embodiments merely represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for processing network model data, characterized in that, The method includes: Receiving a training data acquisition request, where the training data acquisition request carries an identifier of a training data set to be acquired; When the type of the training data set to be acquired corresponding to the identifier of the training data set to be acquired is of text type, searching for a corresponding target text training data set identifier in the first queue; the first queue stores various first text training data set identifiers and corresponding first text training data sets; When the corresponding target text training data set identifier is not found in the first queue, searching for the target text training data set identifier corresponding to the identifier of the training data set to be acquired in the second queue; the second queue stores various second text training data set identifiers, corresponding historical access times, and corresponding disk storage locations, and both the first queue and the second queue are located in the data cache; When the corresponding target text training data set identifier is found in the second queue, obtaining the target disk storage location corresponding to the found target text training data set identifier; Based on the target disk storage location, reading the target text training data set corresponding to the target text training data set identifier from the corresponding disk, updating the historical access time corresponding to the target text training data set identifier in the second queue, and the second text training data set corresponding to the second text training data set identifier in the second queue is used to be written into the first queue when the corresponding historical access time meets the cache condition; Inputting the target text training data set into a network model for training, where the network model is used to process input data according to the model task type.
2. The method according to claim 1, wherein The first queue stores various first text training data set identifiers, corresponding first text training data sets, corresponding cache storage locations, and corresponding first access time points; After searching for the corresponding target text training data set identifier in the first queue based on the identifier of the training data set to be acquired, it further includes: When the corresponding target text training data set identifier is found in the first queue, obtaining the target cache storage location corresponding to the target text training data set identifier, and reading the target text training data set corresponding to the target text training data set identifier from the corresponding cache based on the target cache storage location; Obtaining the current time point, and updating the first access time point corresponding to the target text training data set identifier based on the current time point.
3. The method according to claim 2, wherein The training data acquisition request also carries the data volume of the training data set to be acquired; The step of, when the corresponding target text training data set identifier is found in the first queue, obtaining the target cache storage location corresponding to the target text training data set identifier, and reading the target text training data set corresponding to the target text training data set identifier from the corresponding cache based on the target cache storage location, includes: When the corresponding target text training data set identifier is found in the first queue, obtaining the cache data volume of the target text training data set corresponding to the target text training data set identifier in the first queue; When the amount of cached data exceeds the amount of data of the training data set to be obtained, the target cache storage location corresponding to the target text training data set identifier is obtained, and the target text training data set of the amount of data of the training data set to be obtained is read from the corresponding cache based on the target cache storage location.
4. The method according to claim 3, wherein After obtaining the amount of cached data of the target text training data set corresponding to the target text training data set identifier in the first queue when finding the corresponding target text training data set identifier in the first queue, it further includes: When the amount of cached data does not exceed the amount of data of the training data set to be obtained, the target disk storage location corresponding to the target text training data set identifier is obtained, and the target text training data set of the amount of data of the training data set to be obtained is read from the corresponding disk based on the target disk storage location.
5. The method according to claim 1, wherein After finding the target text training data set identifier corresponding to the training data set to be obtained in the second queue, it further includes: When the target text training data set identifier corresponding to the training data set to be obtained is not found in the second queue, the download address of the training data set to be obtained is obtained based on the training data set to be obtained identifier; Download the target text training data set based on the download address of the training data set to be obtained, store the target text training data set in the disk, and obtain the disk storage address and the current access times; Write the training data set to be obtained identifier, the current access times, and the disk storage address into the second queue.
6. The method according to claim 5, wherein The writing the training data set to be obtained identifier, the current access times, and the disk storage address into the second queue includes: When the second queue is full, find the identifier of the second text training data set to be deleted corresponding to the target time point in the second queue, where the target time point refers to the oldest time point; Delete the identifier of the second text training data set to be deleted corresponding to the target time point, the historical access times to be deleted corresponding to the identifier of the second text training data set to be deleted, and the disk storage address to be deleted from the second queue; Write the training data set to be obtained identifier, the current access times, and the disk storage address into the second queue.
7. The method according to claim 1, wherein The writing into the first queue includes: When the first queue is full, find the identifier of the first text training data set to be deleted corresponding to the target time point in the first queue, where the target time point refers to the oldest time point; Delete the identifier of the first text training data set to be deleted corresponding to the target time point and the first text training data set corresponding to the identifier of the first text training data set to be deleted from the first queue; Obtain the current access time point, and write the current access time point, the target text training data set identifier, and the target text training data set corresponding to the target text training data set identifier into the first queue.
8. The method according to claim 1, wherein After receiving the training data acquisition request, where the training data acquisition request carries the training data set to be obtained identifier, it further includes: When the type of the training dataset to be obtained corresponding to the training dataset identifier to be obtained is of the image type, obtain the corresponding image download address according to the image training dataset identifier, and perform a download based on the image download address to obtain an image training dataset; Input the image training dataset into a network model for training.
9. A data display method, characterized in that, The method includes: Receive a data display instruction sent through a data display page, and send a data acquisition request to a server based on the data display instruction. The data acquisition request carries a dataset identifier to be displayed. When the type of the dataset to be displayed corresponding to the dataset identifier to be displayed is of the text type, the server is configured to find a corresponding target text dataset identifier in a first queue. The first queue stores each first text dataset identifier and the corresponding first text dataset. When the corresponding target text dataset identifier is not found in the first queue, find the target text dataset identifier corresponding to the dataset identifier to be displayed in a second queue. The second queue stores each second text dataset identifier, the corresponding historical access times, and the corresponding disk storage location. Both the first queue and the second queue are located in a data cache. When the corresponding target text dataset identifier is found in the second queue, obtain the target disk storage location corresponding to the found target text dataset identifier, read the target text dataset corresponding to the target text dataset identifier from the corresponding disk based on the target disk storage location, update the historical access times corresponding to the target text dataset identifier in the second queue. The second text dataset corresponding to the second text dataset identifier in the second queue is configured to be written into the first queue when the corresponding historical access times meet the cache condition, and return the target text dataset to the terminal; Obtain the target text dataset returned by the server, and display the target text dataset through the data display page.
10. The method according to claim 9, wherein The method further includes: Receive a data display instruction sent through a data display page, and send a data synchronization request to a server based on the data display instruction. The data synchronization request carries a dataset identifier to be displayed. The server responds to the data synchronization request, generates an asynchronous task identifier, and records that the asynchronous task status corresponding to the asynchronous task identifier is not completed, and returns the asynchronous task identifier to the terminal; The server is used to execute the asynchronous task corresponding to the asynchronous task identifier. The asynchronous task refers to the task of the server to obtain the target text dataset when the type of the dataset to be displayed corresponding to the dataset to be displayed identifier is of text type. The obtaining of the target text dataset means to search for the corresponding text dataset identifier in the first queue by using the dataset to be displayed identifier. When the corresponding text dataset identifier is not found in the first queue, search for the text dataset identifier corresponding to the dataset to be displayed identifier in the second queue. When the corresponding text dataset identifier is found in the second queue, obtain the disk storage location corresponding to the found text dataset identifier, read the text dataset corresponding to the text dataset identifier from the corresponding disk based on the disk storage location, update the historical access times corresponding to the text dataset identifier in the second queue, and update the asynchronous task status corresponding to the asynchronous task identifier to completed; Obtain the asynchronous task identifier returned by the server, and send an asynchronous request to the server at a preset time interval based on the asynchronous task identifier. The server responds to the asynchronous request, queries the asynchronous task status, and when the asynchronous task status is completed, returns the target text dataset corresponding to the dataset to be displayed identifier to the terminal; Obtain the target text dataset returned by the server, and display the target text dataset through the data display page.
11. The method according to claim 9, wherein The method further includes: Receive a data display instruction sent through the data display page, and send a data acquisition request to the server based on the data display instruction. The data acquisition request carries the dataset to be displayed identifier. When the type of the dataset to be displayed corresponding to the dataset to be displayed identifier is of picture type, the server obtains the corresponding picture storage address according to the picture dataset identifier, converts the picture storage address to obtain a picture download address, and returns the picture download address to the terminal; Receive the picture download address returned by the server, and perform picture download based on the picture download address to obtain a picture dataset; Display the picture dataset on the data display page.
12. A network model data processing device, characterized in that, The device includes: A request receiving module, configured to receive a training data acquisition request, where the training data acquisition request carries the identifier of the training dataset to be acquired; A first search module, configured to, when the type of the training dataset to be acquired corresponding to the identifier of the training dataset to be acquired is of text type, search for the corresponding target text training dataset identifier in the first queue; the first queue stores each first text training dataset identifier and the corresponding first text training dataset; A second search module, configured to search for a target text training dataset identifier corresponding to the to-be-acquired training dataset identifier in a second queue when the corresponding target text training dataset identifier cannot be found in the first queue; the second queue stores various second text training dataset identifiers, corresponding historical access times, and corresponding disk storage locations, and both the first queue and the second queue are located in a data cache; A location acquisition module, configured to acquire a target disk storage location corresponding to the found target text training dataset identifier when the corresponding target text training dataset identifier is found in the second queue; A data reading module, configured to read the target text training dataset corresponding to the target text training dataset identifier from the corresponding disk based on the target disk storage location, update the historical access times corresponding to the target text training dataset identifier in the second queue, and the second text training datasets corresponding to the second text training dataset identifiers in the second queue are configured to be written into the first queue when the corresponding historical access times meet the cache conditions; A training module, configured to input the target text training dataset into a network model for training, and the network model is configured to process input data according to the model task type.
13. A data display device, characterized in that, The apparatus includes: A request sending module, configured to receive a data display instruction sent through a data display page, and send a data acquisition request to a server based on the data display instruction, where the data acquisition request carries a to-be-displayed dataset identifier, and the server is configured to search for a corresponding target text dataset identifier in a first queue based on the to-be-displayed dataset identifier when the type of the to-be-displayed dataset corresponding to the to-be-displayed dataset identifier is a text type. The first queue stores various first text dataset identifiers and corresponding first text datasets. When the corresponding target text dataset identifier cannot be found in the first queue, search for the target text dataset identifier corresponding to the to-be-displayed dataset identifier in the second queue. The second queue stores various second text dataset identifiers, corresponding historical access times, and corresponding disk storage locations, and both the first queue and the second queue are located in a data cache. When the corresponding target text dataset identifier is found in the second queue, acquire the target disk storage location corresponding to the found target text dataset identifier, read the target text dataset corresponding to the target text dataset identifier from the corresponding disk based on the target disk storage location, update the historical access times corresponding to the target text dataset identifier in the second queue, and the second text datasets corresponding to the second text dataset identifiers in the second queue are configured to be written into the first queue when the corresponding historical access times meet the cache conditions, and return the target text dataset to the terminal; A text display module, configured to acquire the target text dataset returned by the server and display the target text dataset through the data display page.
14. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.
15. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Parallelized forensic analysis using cloud-based servers
US10740151B1
Apparatuses, methods, and systems for memory interface circuit allocation in a configurable spatial accelerator
US20200310994A1