Data stream analysis method and device and storage medium
By analyzing the log data generated by the streaming media server, the problem of abnormal positioning of the streaming media server is solved, and the automated positioning of abnormal data flow and accurate positioning of problem points is achieved, improving the efficiency of problem investigation and user experience.
Patent Information
- Application Number
- CN202410010350.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-07-04
AI Technical Summary
When the streaming media server works abnormally, it is difficult to accurately locate the specific location where the data flow quality deteriorates, affecting the user experience.
By obtaining the log data generated by the streaming media server, analyzing the log data to determine the business relationship between the server and the data stream, and using data filtering, aggregation, compression, mining and visualization algorithms to process the log data to achieve automatic positioning of abnormal data streams.
It realizes accurate positioning of abnormalities in streaming media servers, quickly identify problem points and displays the information of responsible persons, and improves problem investigation efficiency and user experience.
Smart Images

Figure CN120263629A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of streaming media technology, and more particularly, to a data stream analysis method, apparatus, and storage medium. Background Art
[0002] Streaming media technology can compress continuous multimedia data. After the compressed multimedia data is segmented, multiple data segments can be obtained, and then the server can transmit the multiple data segments to the user sequentially or in real time. The user can watch or listen to the multimedia content while downloading, without having to wait for all the multimedia data to be downloaded to their computer before watching or listening to the multimedia content.
[0003] When the streaming media server malfunctions, the quality of the data stream transmitted by the streaming media server deteriorates, which has an adverse impact on the user experience. Therefore, an effective solution is needed to monitor the working condition of the streaming media server.
[0004] In the related art, the data stream of multimedia may flow through multiple streaming media servers, and the flowing direction is difficult to determine. When the quality of the multimedia data stream deteriorates, it is difficult to locate the specific position where the problem occurs in the data stream. Summary of the Invention
[0005] Embodiments of the present invention provide a data stream analysis method, apparatus, and storage medium.
[0006] The data stream analysis method provided by the embodiments of the present invention includes: obtaining a plurality of first log data, each of the first log data being generated by a corresponding streaming media server for the services performed on the data stream; analyzing the plurality of first log data to determine the plurality of streaming media servers corresponding to the plurality of first log data and the services performed by each streaming media server on the data stream.
[0007] Each streaming media server generates corresponding first log data when performing services on the data stream. The first log data includes a work list of all services performed by a corresponding streaming media server on the data stream. Therefore, it is possible to determine in which server the data stream has worked based on the first log data, and all services performed by each server on the data stream. When troubleshooting the streaming media, the streaming media server with abnormal data stream operation can be accurately located based on the first log data.
[0008] In some embodiments, the data stream analysis method includes: obtaining a plurality of second log data, each of the second log data being generated by a corresponding streaming media server for a corresponding data stream; analyzing the plurality of second log data to determine the plurality of streaming media servers and the plurality of data streams corresponding to the plurality of second log data, and the services executed by each streaming media server for each data stream.
[0009] When troubleshooting data streams, it is possible to accurately locate which data streams are abnormal data streams based on the second log data, where the problem points of the abnormal data streams are specifically on which streaming media server, and each service executed by the streaming media server for the abnormal data stream.
[0010] In some embodiments, after obtaining the plurality of first log data, the data stream analysis method includes: formatting the plurality of first log data to obtain a plurality of single-format log data, each single-format log data corresponding to a streaming media server; analyzing the plurality of single-format log data to determine the plurality of streaming media servers and the plurality of data streams corresponding to the plurality of single-format log data, and the services executed by each streaming media server for each data stream.
[0011] The log collection module formats the collected logs according to the format set by the policy unit. One log analysis module corresponds to one log format. Based on the log format set by the log analysis module, the log formats obtained by the log analysis module are all the same. The log analysis module can avoid the differential problems caused by different log formats when analyzing log data.
[0012] In some embodiments, the analyzing the plurality of first log data includes: according to the analysis requirements of the first log data, filtering the feature quantities of the plurality of first log data by using a data filtering algorithm to obtain a plurality of filtered log data; analyzing the plurality of filtered log data to determine the plurality of streaming media servers and the plurality of data streams corresponding to the plurality of filtered log data, and the services executed by each streaming media server for each data stream.
[0013] The data filtering algorithm can be set according to the services of the streaming media server. Further, the data filtering algorithm can filter out data irrelevant to subsequent modeling or analysis according to the service situation.
[0014] In some embodiments, the analyzing of the plurality of first log data includes: aggregating the plurality of first log data according to the attributes of the first log data by using a data aggregation algorithm to obtain a plurality of aggregated log data, where the attributes of the first log data at least include the time of the first log data; analyzing the plurality of aggregated log data to determine the plurality of streaming media servers and the plurality of data streams corresponding to the plurality of aggregated log data, and the services performed by each streaming media server on each data stream.
[0015] The data aggregation algorithm can integrate a plurality of data based on a certain attribute of the data to obtain a corresponding data set, which is beneficial to eliminating the differences between different data and enabling the unified analysis of a plurality of data within a data set.
[0016] In some embodiments, the analyzing of the first log data includes: compressing the plurality of first log data by using a data compression algorithm to obtain a plurality of compressed log data, where the storage space occupied by the compressed log data is smaller than the storage space occupied by the first log data.
[0017] The log analysis module can use a data compression algorithm to compress, archive, and deduplicate log data to save storage space and improve data query efficiency.
[0018] In some embodiments, the analyzing of the plurality of first log data includes: mining abnormal feature quantities of the first log data by using a data mining algorithm to obtain the mined log data; analyzing the plurality of mined log data to determine the plurality of streaming media servers and the plurality of data streams corresponding to the plurality of mined log data, and the services performed by each streaming media server on each data stream.
[0019] The log analysis module can perform automated processing and analysis on log data based on a deep learning model to mine potential abnormal events and fault causes in the data streams within the streaming media server.
[0020] In some embodiments, after the analyzing of the plurality of first log data, the data stream analysis method includes: analyzing the plurality of first log data to obtain the analyzed data; visualizing the analyzed data by using a data visualization algorithm to obtain visualized data, where the visualized data at least includes a table or an image of the analyzed log data.
[0021] The data visualization algorithm is used to display the processed log data in the form of charts, reports, etc., so that users can more intuitively observe the processed log data.
[0022] In some embodiments, after obtaining a plurality of first log data, the data stream analysis method includes: cutting each of the first log data according to the time sequence to obtain a plurality of log data slices; analyzing the plurality of log data slices to determine a streaming media server and a data stream corresponding to the plurality of log data slices, and the service performed by the streaming media server on the data stream.
[0023] The log analysis module preferentially reads the cache database to obtain the data of the current day, which is beneficial to the log analysis module quickly reading the log data for analysis and can quickly obtain the analysis result.
[0024] In some embodiments, after analyzing the plurality of first log data, the data stream analysis method includes: in response to an access signal received by an application programming interface, displaying the plurality of streaming media servers corresponding to the plurality of first log data and the service performed by each streaming media server on each data stream.
[0025] After analyzing the first log data, the service content performed by each streaming media server on each data stream determined by the analysis can be provided for use by the portal website of the streaming media system or third-party service docking, so that users can accurately analyze the data stream of the streaming media.
[0026] In some embodiments, after obtaining a plurality of first log data, the data stream analysis method includes: storing the plurality of first log data in the cache database; in the case where the time when the first log data is stored in the cache database is greater than the aging time of the first log data, the first log data is removed from the cache database; analyzing the plurality of first log data in the cache database to determine a plurality of streaming media servers corresponding to the plurality of first log data and the service performed by each streaming media server on the data stream.
[0027] An embodiment of the present invention provides a data stream analysis device for a streaming media server. The data stream analysis device includes a log collection module and a log analysis module. Among them, the log collection module is used to obtain a plurality of first log data, and each first log data is generated by the service performed by a corresponding streaming media server on a data stream. The log analysis module is used to analyze the plurality of first log data to determine a plurality of streaming media servers corresponding to the plurality of first log data and the service performed by each streaming media server on the data stream.
[0028] In some embodiments, the data stream analysis device includes a front-end display module. The front-end display module is used to display a plurality of the streaming media servers corresponding to the plurality of the first log data and the services performed on the data stream by each of the streaming media servers.
[0029] An embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium is used to store a computer program, and when the computer-readable storage medium is executed, it implements the data stream analysis method of any one of the above embodiments.
[0030] An embodiment of the present invention provides a data stream analysis method, device and storage medium. The data stream analysis method includes: obtaining a plurality of first log data, each of the first log data being generated by a service performed on the data stream by a corresponding streaming media server; analyzing the plurality of first log data to determine a plurality of streaming media servers corresponding to the plurality of first log data and the services performed on the data stream by each of the streaming media servers. Each streaming media server generates corresponding first log data when performing services on the data stream. The first log data includes a work list of all services performed on the data stream by a corresponding streaming media server. Therefore, it is possible to determine in which server the data stream has worked based on the first log data, and all services performed on the data stream by each server. When troubleshooting problems in the streaming media, the streaming media server with abnormal data stream operation can be accurately located based on the first log data.
[0031] When troubleshooting problems with the data stream, the problem point of the abnormal data stream can be accurately located on which streaming media server and each service performed on the abnormal data stream by the streaming media server based on the first log data. After determining the specific location of the problem point of the abnormal data stream, it is possible to further determine which maintenance personnel or operation and maintenance personnel need to be found for the specific location of the problem point, and then determine who is responsible for the operation and maintenance and development of the streaming media server with problems, so that it is possible to achieve automatic problem location and display the responsible person information.
[0032] The additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0033] The above and / or additional aspects and advantages of the present invention will become apparent and be easily understood from the description of the embodiments in conjunction with the following drawings, where:
[0034] Figure 1 is a flowchart of the data stream analysis method according to the first embodiment of the present invention;
[0035] Figure 2It is a schematic diagram of the data flow analysis device according to an embodiment of the present invention;
[0036] Figure 3 It is a schematic diagram of the flow direction of the data flow according to an embodiment of the present invention;
[0037] Figure 4 It is a schematic diagram of the business process executed by the data flow analysis device according to an embodiment of the present invention;
[0038] Figure 5 It is a schematic diagram of the tree structure of the flow call portrait according to an embodiment of the present invention;
[0039] Figure 6 It is a schematic diagram of the process of the data flow analysis method according to the second embodiment;
[0040] Figure 7 It is a schematic diagram of the circular queue for storing log data according to an embodiment of the present invention;
[0041] Figure 8 It is a schematic diagram of the star structure of the log data according to an embodiment of the present invention. Detailed Embodiment
[0042] The following details the embodiments of the present invention. The described embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0043] Streaming media technology can compress continuous multimedia data. The compressed multimedia data can be cut into multiple data segments, and then the server can transmit the multiple data segments to the user sequentially or in real time in segments. The user can watch and listen to the multimedia content while downloading, without having to wait until all the multimedia data is downloaded to their computer before watching and listening to the multimedia content.
[0044] When the streaming media server malfunctions, the quality of the data stream transmitted by the streaming media server deteriorates, which has an adverse impact on the user experience. Therefore, an effective solution is needed to monitor the working condition of the streaming media server.
[0045] In the related art, the data stream of multimedia may flow through multiple streaming media servers, and the flowing direction is difficult to determine. When the quality of the multimedia data stream deteriorates, it is difficult to locate the specific position where the problem occurs in the data stream.
[0046] Refer to Figure 1 , an embodiment of the present invention provides a data flow analysis method. In some embodiments, the data flow analysis method includes:
[0047] Step S100: Obtain multiple first log data, where each first log data is generated by a corresponding streaming media server for the services executed on the data stream.
[0048] Step S200: Analyze the multiple first log data to determine the multiple streaming media servers corresponding to the multiple first log data and the services executed by each streaming media server on the data stream.
[0049] Each streaming media server will generate corresponding first log data when executing services on the data stream. The first log data includes a work list of all services executed by a corresponding streaming media server on the data stream. Therefore, it is possible to determine in which server the data stream has worked based on the first log data, as well as all the services executed by each server on the data stream. When troubleshooting problems in the streaming media, the streaming media server with abnormal data stream operation can be accurately located based on the first log data.
[0050] Specifically, in step S100, each streaming media server will generate corresponding first log data when executing services on the data stream. The first log data can be stored in the log file corresponding to the streaming media server. One first log data includes a list of all services executed by a corresponding streaming media server.
[0051] In step S200, based on the obtained first log data, it is possible to determine the services executed by each streaming media server on the data stream, and further determine the working status of the data stream in each streaming media.
[0052] When troubleshooting problems with the data stream, the problem point of the abnormal data stream can be accurately located on which streaming media server and each service executed by the streaming media server on the abnormal data stream based on the first log data.
[0053] After determining the specific location of the problem point of the abnormal data stream, it is possible to further determine which maintenance personnel or operation and maintenance staff need to be found for the specific location of the problem point, and then determine who is responsible for the operation and maintenance and development of the streaming media server with problems, so as to achieve automatic problem location and display of responsible person information.
[0054] Refer to Figure 2 , an embodiment of the present invention also provides a data stream analysis device 100. In some embodiments, the data stream analysis method can be implemented by the data stream analysis device 100, that is to say, the data stream analysis device 100 is used to implement the control method.
[0055] Of course, in other embodiments, the control method may also be implemented by other devices or apparatuses, and is not limited to being implemented by the data flow analysis apparatus 100. The data flow analysis apparatus 100 may also not be dedicated to implementing the control method of the embodiments of the present invention, but may implement other functions and methods.
[0056] In some embodiments, the data flow analysis apparatus 100 includes a log collection module 10 and a log analysis module 20. Among them, the log collection module 10 is used to obtain a plurality of first log data, and each first log data is generated by a corresponding streaming media server for the service executed on the data stream. The log analysis module 20 is used to analyze the plurality of first log data to determine the plurality of streaming media servers corresponding to the plurality of first log data and the service executed by each streaming media server on the data stream.
[0057] If additional service monitoring of the data stream is added to the service layer inside the streaming media server, it will affect the execution logic of the original main service of the streaming media server. The data flow analysis apparatus 100 monitors the service executed by each streaming media server on the data stream by analyzing the first log data generated by the streaming media server, making the data flow analysis method more professional and independent. By setting up an additional data flow analysis apparatus 100 to analyze the service executed by the streaming media server on the data stream, it will not affect the execution logic of the original main service of the server, enabling the streaming media server to have a free architecture to build and expand new services.
[0058] The data flow analysis apparatus 100 can understand the basic overview of the data stream by analyzing the first log data. For example: where is the problem point of the abnormal data stream, who can solve this problem, the statistical data of the abnormal data stream, and the portrait of the internal call of the data stream. The above basic overview of the data stream can provide necessary data support for user decision-making.
[0059] Refer to Figure 3 and Figure 4 , in some embodiments, after analyzing the plurality of first log data, the data flow analysis method includes: in response to an access signal received by the application programming interface, displaying the plurality of streaming media servers corresponding to the plurality of first log data and the service executed by each streaming media server on the data stream.
[0060] After analyzing the first log data, the service content executed by each streaming media server on each data stream determined by the analysis can be provided for use by the portal website of the streaming media system or for third-party service docking, so that users can accurately analyze the data stream of the streaming media.
[0061] Specifically, an Application Programming Interface (API) can be a set of predefined functions. Users or third-party services can access the API interface through simple instructions to obtain the service content executed by each streaming media server for the data stream.
[0062] The service content executed by each streaming media server for the data stream can be called to the portal website (Portal) of the streaming media system or a third-party service. Users can obtain the collected first log data and the results obtained by analyzing the first log data through the Portal or the user interface of the third party, and understand the running status of the background streaming server.
[0063] Furthermore, a corresponding API interface can be set for the abnormal data stream, which is convenient for users to quickly view the detailed working conditions of the abnormal data stream. For example, where is the abnormal point of the data stream, which maintenance personnel or operation and maintenance personnel need to solve the abnormal point of the data stream, and then determine which specific operation and maintenance personnel and developers are responsible for the streaming media server with problems, so as to achieve automatic problem location and display of the responsible person information.
[0064] In some embodiments, the data stream analysis device 100 includes a front-end display module, which is used to display multiple streaming media servers corresponding to multiple first log data and the services executed by each streaming media server for the data stream.
[0065] After analyzing the first log data, the front-end display module can display the service content of each streaming media server pair. The processed first log data can be displayed in the form of charts, reports, etc., so that users can more intuitively understand the working status of each streaming media server.
[0066] Specifically, the front-end display module can be the display module of the user interface. Users can intuitively understand the working status of each streaming media server through the Portal or the user interface of the third party.
[0067] The working status of each streaming media server can be displayed in the form of charts, reports, etc., so that users can more intuitively understand the working status of each streaming media server. The abnormal data stream can be displayed in a specific color, which is convenient for users to quickly view the detailed working conditions of the abnormal data stream.
[0068] Users can quickly query the information corresponding to the data stream by accessing the API interface through simple instructions, and the front-end display module displays the working conditions corresponding to the data stream.
[0069] For example, when the user needs to observe the location of the anomaly point in the data stream, the user can send a command to query the location of the anomaly point in the data stream on the interaction interface. At this time, the front-end display module can display where the anomaly point in the data stream is, which maintenance personnel or operation and maintenance personnel need to solve the anomaly point in the data stream, and specific information about which streaming media server is responsible for operation and maintenance and development when problems occur.
[0070] Refer to Figure 5 , after the log analysis module 20 completes the analysis of the first log data, the data of the log analysis module 20 is called through the API interface, and the display module can display the corresponding stream call portrait. The main attributes of the stream call portrait include: service label (ID), stream label (ID), entry time, departure time, status code, error message, and reporting time.
[0071] Specifically, among the main attributes of the stream call portrait, each streaming media server has a corresponding service ID, and the corresponding streaming media server can be determined according to the service ID. The service ID can be used as an isolation parameter for the streaming media server to execute services, and each streaming media server is only responsible for its own business layer. When the data flows to other streaming media servers, the corresponding business is no longer managed by the previous streaming media server.
[0072] For example, the service ID corresponding to the streaming media server A can be A, and the service ID corresponding to the streaming media server B can be B. The stream call portrait with the service ID of A is obtained based on the log of the streaming media server A, and the stream call portrait with the service ID of B is obtained based on the log of the streaming media server B.
[0073] Since a first log data is generated based on a corresponding streaming media server, therefore, the service ID of the stream call portrait corresponding to a first log data is the same.
[0074] Furthermore, among the main attributes of the stream call portrait, each data stream has a corresponding stream ID, and the corresponding data stream can be determined according to the stream ID. Each data stream has a unique stream ID from entering the system to ending, and the stream ID will not change within different streaming media servers and can be used as a marker to track and identify the data stream.
[0075] A data stream can flow through multiple streaming media servers. The stream ID of the stream call portrait corresponding to a data stream remains unchanged, and the service ID changes according to the corresponding streaming media server. Multiple data streams can exist within a streaming media server. The service ID of the stream call portrait corresponding to a streaming media server remains unchanged, and the stream ID changes according to the corresponding data stream.
[0076] Refer to Figure 6 , in some embodiments, the data stream analysis method includes:
[0077] Step S110: Obtain multiple second log data, where each second log data is generated by a corresponding streaming media server for a corresponding data stream during the execution of operations.
[0078] Step S210: Analyze the multiple second log data to determine the multiple streaming media servers and multiple data streams corresponding to the multiple second log data, as well as the operations performed by each streaming media server on each data stream.
[0079] When troubleshooting data streams, it is possible to accurately locate which data streams are abnormal data streams based on the second log data, where the problem points of the abnormal data streams are specifically on which streaming media server, and each operation performed by the streaming media server on the abnormal data stream.
[0080] Specifically, each streaming media server generates corresponding second log data when performing operations on each data stream. The second log data includes a work list of all operations performed by a corresponding streaming media server on a corresponding data stream. Therefore, it is possible to determine on which server each data stream has specifically worked based on the second log data, as well as all operations performed by each server on each data stream.
[0081] After determining which data streams are abnormal data streams, it is possible to further determine which maintenance personnel or operation and maintenance personnel need to be sought for the specific location of the abnormal data streams, and then determine who is specifically responsible for the operation and maintenance and development of the streaming media server where the problem occurs, so as to achieve automatic problem location and display of responsible person information.
[0082] When a streaming media server performs operations on a data stream, it can generate a second log data. Therefore, the service ID and stream ID of the stream call portrait corresponding to a second log data are the same.
[0083] In some embodiments, after obtaining multiple first log data, the data stream analysis method includes: formatting the multiple first log data to obtain multiple single-format log data, where each single-format log data corresponds to a streaming media server; analyzing the multiple single-format log data to determine the multiple streaming media servers and multiple data streams corresponding to the multiple single-format log data, as well as the operations performed by each streaming media server on each data stream.
[0084] The log collection module 10 can collect the log files generated by the streaming media server according to a predefined collection format. The log collection module 10 formats them and sends them to the log analysis module 20. The first log data corresponding to a streaming media server can report the collected log data to the log analysis module 20 in a unified format.
[0085] Specifically, the log analysis module 20 includes a policy unit, a collection unit, and a management unit. The policy unit, the collection unit, and the management unit can issue configuration instructions to the log analysis module 20, and the log collection module 10 can also work according to the configuration instructions issued by the log analysis module 20.
[0086] Among them, the management unit can set the speed at which the log collection module 10 collects log data. The collection unit of the log analysis module 20 is used to interface with the log collection module 10. One log analysis module 20 can interface with multiple log collection modules 10, and the format of the log data uploaded by each log collection module 10 to the log analysis module 20 is the same.
[0087] The policy unit can set the format of the log, and the log collection module 10 formats the collected log according to the format set by the policy unit. One log analysis module 20 corresponds to one log format. Based on the log format set by the log analysis module 20, the log formats obtained by the log analysis module 20 are the same, and the log analysis module 20 can avoid the differential problems caused by different log formats when analyzing log data.
[0088] Each streaming media server can correspond to one log analysis module 20, and each log analysis module 20 sets a log format. The first log data corresponding to one streaming media server can report the collected log data to the log analysis module 20 in a unified format.
[0089] Before the log collection module 10 formats the log data, the log collection module 10 must first read the log data and check whether its format is in the list of supported formats of the log collection module 10. If the log collection module 10 cannot support the format of the read log data, the log collection module 10 fails to collect and issues a prompt signal for the error of collecting log data. If the log collection module 10 can support the format of the read log data, the log collection module 10 can format the collected log data according to the agreed format and report it to the log analysis module 20.
[0090] In some embodiments, the log format set by the log analysis module 20 can be in the K-V format, where one K corresponds to one V and is connected by an equal sign, and multiple K-Vs are separated by spaces. For example, the K-V format can be time="2023-04-18 09:13:31.623455" level=debug model="streaming-server / pull / rtsp / client" sid=001722053719.
[0091] Common Ks include time, level, msg, model, etc. For some important business process logs, some key Ks can also be included, such as: sid, tid, code, addr, etc. Users can also customize K according to the actual business needs.
[0092] In some embodiments, after obtaining a plurality of first log data, the data stream analysis method includes: cutting each first log data according to the time sequence to obtain a plurality of log data slices; analyzing the plurality of log data slices to determine a streaming media server and a data stream corresponding to the plurality of log data slices, and the service performed by the streaming media server on the data stream.
[0093] Specifically, after the log collection module 10 collects the log data, the log collection module 10 can cut the collected log data according to the time sequence, and distribute the log data slices to the log analysis module 20 according to the time sequence. The log analysis module 20 analyzes the log data slices according to the corresponding algorithm, and outputs the analyzed result data for the corresponding API interface to call.
[0094] In some embodiments, the log analysis module 20 further includes a statistics unit and an analysis unit. The statistics unit can count the number of parallel tasks, the number of abnormal tasks, the number and frequency of each error occurrence, the total number of task flow logs, the total number of logs per hour and per day, the total number of access collection services, the number of different types of logs, the log flow rate of different types, etc. on each streaming media server. The analysis data background independent thread performs operations and stores the results in the cache service for the API interface to query and use.
[0095] The policy unit can also set the working mode of the log analysis module 20. The analysis module can analyze the log data according to the set working mode to obtain the corresponding analysis results. The analysis result data obtained by the analysis module includes: single dimension, multi-dimension, keyword, time period, fuzzy query, error rate, success rate, time series data, portrait data, etc.
[0096] In some embodiments, analyzing a plurality of first log data includes: filtering the feature quantities of the plurality of first log data according to the analysis requirements of the first log data by using a data filtering algorithm to obtain a plurality of filtered log data; analyzing the plurality of filtered log data to determine a plurality of streaming media servers and a plurality of data streams corresponding to the plurality of filtered log data, and the service performed by each streaming media server on each data stream.
[0097] The data filtering algorithm can be set according to the service of the streaming media server. Further, the data filtering algorithm can filter out the data irrelevant to subsequent modeling or analysis according to the service situation.
[0098] Specifically, according to the business conditions of the streaming media server, specific eigenvalue screening is performed, screening conditions for specific eigenvalues are specified, and then data that does not meet the corresponding screening conditions is filtered. For example, it can be specified that a certain eigenvalue needs to be greater than a certain value or less than a certain value, and the data filtering algorithm can filter the data that does not meet the corresponding screening conditions.
[0099] In some embodiments, analyzing multiple first log data includes: aggregating multiple first log data according to the attributes of the first log data by using a data aggregation algorithm to obtain multiple aggregated log data, where the attributes of the first log data at least include the time of the first log data; analyzing the multiple aggregated log data to determine multiple streaming media servers, multiple data streams corresponding to the multiple aggregated log data, and the services performed by each streaming media server on each data stream.
[0100] The data aggregation algorithm can integrate multiple data based on a certain attribute of the data to obtain a corresponding data set, which is beneficial to eliminating the differences between different data and enabling unified analysis of multiple data within a data set.
[0101] Specifically, the log data can be aggregated according to the time of obtaining the log data. For example, all log data obtained within a certain period of time can be aggregated into a corresponding data set, and the log analysis module 20 can analyze this data set to determine the services performed by the streaming media server on the data stream during this period of time.
[0102] The log data can be aggregated according to the service ID corresponding to the log data, that is, all log data corresponding to a streaming media server is used as a data set, and the log analysis module 20 can analyze this data set to determine the services performed by this streaming media server on the data stream.
[0103] The log data can be aggregated according to the stream ID corresponding to the log data, that is, all log data corresponding to a data stream is used as a data set, and the log analysis module 20 can analyze this data set to determine the services performed by this data stream.
[0104] In some embodiments, analyzing the first log data includes: compressing multiple first log data by using a data compression algorithm to obtain multiple compressed log data, where the storage space occupied by the compressed log data is smaller than the storage space occupied by the first log data.
[0105] The log analysis module 20 can use the data compression algorithm to compress, archive, and deduplicate the log data to save storage space and improve data query efficiency.
[0106] Specifically, the data flow analysis device 100 further includes a storage area, and the log data collected by the log collection module 10 and the results determined by the log analysis module 20 are stored in the corresponding storage areas. Users can query the log data and analysis results in the storage area through the API interface.
[0107] The data compression algorithm can remove the redundancy of the log data and reduce the storage space occupied by the log data. When the user queries the log data and analysis results in the storage area through the API interface, since the amount of log data becomes smaller, the user can query the data to be queried more quickly.
[0108] In some embodiments, analyzing multiple first log data includes: mining abnormal feature quantities of the first log data using a data mining algorithm to obtain mined log data; analyzing multiple mined log data to determine multiple streaming media servers and multiple data flows corresponding to the multiple mined log data, and the services performed by each streaming media server on each data flow.
[0109] The log analysis module 20 can perform automated processing and analysis on the log data based on a deep learning model to mine potential abnormal events and failure causes in the data flow within the streaming media server.
[0110] Specifically, deep learning is a machine learning algorithm based on neural networks. The data mining algorithm based on deep learning can include autoencoders, convolutional neural networks, recurrent neural networks, etc. The features of the log data can be extracted based on the deep learning model to mine potential abnormal events and failure causes in the data flow within the streaming media server.
[0111] An autoencoder is an unsupervised learning algorithm whose purpose is to learn the process of converting input data into feature values. An autoencoder includes an encoder and a decoder. The encoder maps the input data to the feature value space, and the decoder maps these feature values back to the original data space. During the training process, the autoencoder learns the representation of the feature values by minimizing the reconstruction error. The obtained feature values can be used as the input for subsequent data mining tasks.
[0112] A convolutional neural network consists of multiple convolutional layers and pooling layers. The convolutional layers are responsible for extracting local features in the image, and the pooling layers are used to reduce the spatial dimension of the features. The convolutional neural network can automatically learn the features in the image, saving the process of manually designing features.
[0113] In some embodiments, after analyzing multiple first log data, the data flow analysis method includes: analyzing multiple first log data to obtain analyzed data; visualizing the analyzed data using a data visualization algorithm to obtain visualized data, where the visualized data at least includes a table or an image of the analyzed log data.
[0114] The processed log data is presented in the form of charts, reports, etc. by using data visualization algorithms, so that users can more intuitively observe the processed log data.
[0115] Specifically, users can query the corresponding statistical analysis charts and anomaly analysis charts through the API interface. The log analysis module 20 can present the statistical analysis data in the form of statistical analysis charts, and the log analysis module 20 can present the log data of identified work anomalies in the form of log analysis charts.
[0116] In some embodiments, after analyzing multiple first log data, the data stream analysis method includes: storing the first log data collected by the log collection module 10 and the results determined by the log analysis module 20.
[0117] In some embodiments, the data stream analysis device 100 further includes a storage area, and the first log data collected by the log collection module 10 and the results determined by the log analysis module 20 are stored in the corresponding storage area.
[0118] In some embodiments, after obtaining multiple first log data, the data stream analysis method includes: storing the multiple first log data in the cache database; in the case where the time when the first log data is stored in the cache database is greater than the aging time of the first log data, the first log data is removed from the cache database; analyzing the multiple first log data in the cache database to determine the multiple streaming media servers corresponding to the multiple first log data and the services performed by each streaming media server on the data stream.
[0119] Specifically, the data stream analysis device 100 may further include a corresponding database, which includes a cache database and a big data storage database. Among them, the cache database can cache the log data of the current day, and can also cache the data for the operation of the log analysis module 20. For example, the statistical analysis results and configuration data of the log analysis module 20.
[0120] The log analysis module 20 preferentially reads the cache database to obtain the data of the current day, which is beneficial for the log analysis module 20 to quickly read and analyze the log data and can quickly obtain the analysis results.
[0121] When the amount of log data of the current day is small, the cache database can cache all the log data of the current day. When the amount of log data of the current day is large enough, the cache database stops caching after the cached log data reaches a certain quantity.
[0122] When the data in the cache database exceeds a certain quantity or the storage time exceeds a predetermined time, the data is stored in the big data storage database. The big data storage database can serve as a historical database to backup the historical log data. When the user needs to analyze the historical log data, they can read the big data storage database to obtain the historical log data.
[0123] Refer to Figure 7 , in some embodiments, the log collection module 10 designs a circular queue to store the collected log data. The positions in the circular queue can be sorted. The newly obtained log data is written to the entrance of the circular queue, that is, the first position of the circular queue. When the log data is written to the entrance of the circular queue, the positions of all the log data in the circular queue are shifted up by one position. When the position of the log data exceeds the upper limit of the circular queue, the log data is stored in the big data storage database as historical log data or discarded.
[0124] For example, the length of the circular queue can be 100,000 entries, and each entry can store one log data. The first position of the circular queue is the entrance, and the 100,000th position of the circular queue is the exit. Each time the log collection module 10 collects a log data, the positions of all the log data in the circular queue are shifted up by one position. When the position of the log data exceeds 100,000, the log data is stored in the big data storage database as historical log data or discarded.
[0125] Refer to Figure 8 , the analysis result finally output by the log analysis module 20 is time-series linear star-structured log data. The star-structured data includes all the data at that time point. According to the star structure of the log data, the services and working conditions performed by the streaming media server on each data stream at different time points can be determined.
[0126] The embodiments of the present invention provide a computer-readable storage medium. The computer-readable storage medium is used to store a computer program, and when the computer-readable storage medium is executed, it implements the data stream analysis method of any of the above embodiments.
[0127] The embodiments of the data stream analysis method can refer to the above embodiments. The beneficial effects of the computer-readable storage medium include all the beneficial effects of the data stream analysis method, which will not be elaborated here one by one.
[0128] An embodiment of the present invention provides a data stream analysis method, apparatus, and storage medium. The data stream analysis method includes: obtaining a plurality of first log data, where each first log data is generated by a corresponding streaming media server for the services executed on the data stream; analyzing the plurality of first log data to determine the plurality of streaming media servers corresponding to the plurality of first log data and the services executed by each streaming media server on the data stream. Each streaming media server generates corresponding first log data when executing services on the data stream. The first log data includes a work list of all services executed by a corresponding streaming media server on the data stream. Therefore, it is possible to determine in which server the data stream has specifically worked and all services executed by each server on the data stream based on the first log data. When troubleshooting problems with the streaming media, the streaming media server with abnormal data stream operation can be accurately located based on the first log data.
[0129] When troubleshooting problems with the data stream, it is possible to accurately locate on which streaming media server the problem point of the abnormal data stream is and each service executed by the streaming media server on the abnormal data stream based on the first log data. After determining the specific location of the problem point of the abnormal data stream, it is possible to further determine which maintenance personnel or operation and maintenance personnel should be sought for the specific location of the problem point, and then determine who is responsible for the operation and development of the streaming media server with the problem, so as to achieve automatic problem location and display of the responsible person information.
[0130] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0131] In addition, the term "connection" should be understood in a broad sense. For example, it may include fixed connection, may also include detachable connection, or integral connection; it may include direct connection, may also be indirectly connected through an intermediate medium, and may also include the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0132] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0133] Any process or method description shown in a flowchart or otherwise described herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations, where functions may be performed in a substantially simultaneous manner or in an order opposite to that shown or discussed, according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0134] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A data flow analysis method, characterized in that, The data stream analysis method includes: Obtain a plurality of first log data, where each of the first log data is generated by a corresponding streaming media server for the services executed on the data stream; Analyze the plurality of first log data to determine the plurality of streaming media servers corresponding to the plurality of first log data and the services executed by each of the streaming media servers on the data stream.
2. The data flow analysis method according to claim 1, wherein The data stream analysis method includes: Obtain a plurality of second log data, where each of the second log data is generated by a corresponding streaming media server for the services executed on a corresponding data stream; Analyze the plurality of second log data to determine the plurality of streaming media servers, the plurality of data streams corresponding to the plurality of second log data, and the services executed by each of the streaming media servers on each of the data streams.
3. The data flow analysis method according to claim 1, characterized in that After obtaining the plurality of first log data, the data stream analysis method includes: Format the plurality of first log data to obtain a plurality of single-format log data, where each single-format log data corresponds to a streaming media server; Analyze the plurality of single-format log data to determine the plurality of streaming media servers, the plurality of data streams corresponding to the plurality of single-format log data, and the services executed by each of the streaming media servers on each of the data streams.
4. The data flow analysis method according to claim 1, wherein The analysis of the plurality of first log data includes: According to the analysis requirements of the first log data, use a data filtering algorithm to filter the feature quantities of the plurality of first log data to obtain a plurality of filtered log data; Analyze the plurality of filtered log data to determine the plurality of streaming media servers, the plurality of data streams corresponding to the plurality of filtered log data, and the services executed by each of the streaming media servers on each of the data streams.
5. The data flow analysis method according to claim 1, wherein The analysis of the plurality of first log data includes: According to the attributes of the first log data, use a data aggregation algorithm to aggregate the plurality of first log data to obtain a plurality of aggregated log data, where the attributes of the first log data at least include the time of the first log data; Analyze the plurality of aggregated log data to determine the plurality of streaming media servers, the plurality of data streams corresponding to the plurality of aggregated log data, and the services executed by each of the streaming media servers on each of the data streams.
6. The data flow analysis method according to claim 1, wherein The analysis of the first log data includes: Use a data compression algorithm to compress the plurality of first log data to obtain a plurality of compressed log data, where the storage space occupied by the compressed log data is smaller than the storage space occupied by the first log data.
7. The data flow analysis method according to claim 1, wherein The analysis of the plurality of first log data includes: Use a data mining algorithm to mine the abnormal feature quantities of the first log data to obtain the mined log data; Analyze the plurality of mined log data to determine the plurality of streaming media servers, the plurality of data streams corresponding to the plurality of mined log data, and the services executed by each of the streaming media servers on each of the data streams.
8. The data flow analysis method according to claim 1, wherein After analyzing the plurality of first log data, the data stream analysis method includes: Analyze the multiple first log data to obtain the analyzed data; Use a data visualization algorithm to visualize the analyzed data to obtain visualized data, where the visualized data at least includes a table or an image of the analyzed log data.
9. The data flow analysis method according to claim 1, wherein After obtaining the multiple first log data, the data stream analysis method includes: Cut each first log data according to the time sequence to obtain multiple log data slices; Analyze the multiple log data slices to determine a corresponding streaming media server and a data stream of the multiple log data slices, and the service performed by the streaming media server on the data stream.
10. The data flow analysis method according to claim 1, characterized in that After analyzing the multiple first log data, the data stream analysis method includes: In response to an access signal received by the application programming interface, display the multiple streaming media servers corresponding to the multiple first log data and the service performed by each streaming media server on each data stream.
11. The data flow analysis method according to claim 1, wherein After obtaining the multiple first log data, the data stream analysis method includes: Store the multiple first log data inside a cache database; In the case where the time when the first log data is stored inside the cache database is greater than the aging time of the first log data, the first log data is removed from the cache database; Analyze the multiple first log data inside the cache database to determine the multiple streaming media servers corresponding to the multiple first log data and the service performed by each streaming media server on the data stream.
12. A data stream analysis device for a streaming media server, characterized in that, The data stream analysis device includes: A log collection module, which is used to obtain multiple first log data, and each first log data is generated by the service performed by a corresponding streaming media server on a data stream; A log analysis module, which is used to analyze the multiple first log data to determine the multiple streaming media servers corresponding to the multiple first log data and the service performed by each streaming media server on the data stream.
13. The data flow analysis device according to claim 12, characterized in that The data stream analysis device includes a front-end display module, which is used to display the multiple streaming media servers corresponding to the multiple first log data and the service performed by each streaming media server on the data stream.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and when the computer-readable storage medium is executed, it implements the data stream analysis method according to any one of claims 1-11.