Media data management method and system based on big data identification
By clustering and simplifying media data through big data identification technology, and dynamically adjusting the data volume according to the performance of the receiving device, the problem of ineffective resource consumption in media data transmission is solved, and efficient resource utilization and effective display of media data are achieved.
Patent Information
- Application Number
- CN202510732118.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-12
AI Technical Summary
In the prior art, a large amount of ineffective network resources are consumed during the process of acquiring media data. In particular, when the device performance is insufficient, large-capacity media data cannot be effectively displayed.
Through a media data management method based on big data identification, including data identification extraction, similarity calculation, clustering, device parameter adjustment and data volume simplification, the transmission volume of media data is dynamically adjusted to reduce invalid resource consumption.
It effectively reduces the amount of invalid network resources during media data transmission, improves resource utilization, and ensures that media data can be effectively displayed on the receiving device.
Smart Images

Figure CN120632157A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management, and in particular to a media data management method and system based on big data identification. Background Art
[0002] With the popularization of smart devices, the amount of media data content is increasing. In particular, with the popularity of video platforms, the proportion of video data in media data is increasing. Correspondingly, the requirements for network resources are also getting higher and higher. Although there are many services with free or unlimited traffic, the network resources consumed are fixed. In fact, many devices cannot display media data with large content volumes. For example, the performance of the display module is insufficient, and it cannot display the original effect of the media data. However, some of the content that cannot be displayed is also sent to the device, which is invalid network resources. Therefore, how to reduce the amount of invalid resources in the process of media data acquisition is the technical problem that the technical solution of the present invention aims to solve. Summary of the Invention
[0003] The purpose of the present invention is to provide a media data management method and system based on big data identification to solve the problems raised in the above background technology.
[0004] To achieve the above object, the present invention provides the following technical solutions:
[0005] A media data management method based on big data identification, the method comprising:
[0006] Acquire media data, identify the media data, and extract a data identifier; wherein the data identifier is extracted in a random manner;
[0007] Comparing data identifiers of the media data, calculating identifier similarity, and clustering the media data based on the identifier similarity;
[0008] For each type of media data, regularly obtain the device parameters of the receiver within a preset time period, and adjust the data volume of the media data of this type according to the device parameters;
[0009] The media data is simplified according to the data volume, and when a new query request is received, the simplified media data is sent to the user.
[0010] As a further solution of the present invention, the steps of acquiring media data, identifying the media data, and extracting the data identifier include:
[0011] Reading media data in sequence from a database, converting the media data to obtain a data sequence;
[0012] Compare adjacent data in the data sequence and calculate data similarity;
[0013] When the data similarity of adjacent data is greater than the preset similarity threshold, one data is randomly removed;
[0014] The loop is executed until the similarity of all adjacent data in the data sequence is less than the preset similarity threshold;
[0015] Identify each data in the data sequence and extract the data identifier.
[0016] As a further solution of the present invention: the step of identifying each data in the data sequence and extracting the data identifier includes:
[0017] Randomly select data from the data sequence as the data to be identified;
[0018] Input the data to be identified into the preset recognition model to extract keywords;
[0019] Count the extracted keywords as data identifiers and synchronously update the selection probability of each data in the data sequence;
[0020] The loop is executed until a preset loop exit condition is met; the loop exit condition is a keyword condition;
[0021] The process of determining the selection probability is as follows:
[0022]
[0023] Where, P i is the probability of selecting the i-th data, T i is the characteristic value of the i-th data, N is the total number of data, M is the number of selected data, and σ is the preset parameter; Z j is the calibration value of the j-th selected data, and the calibration value is a preset value.
[0024] As a further solution of the present invention, the steps of comparing data identifiers of media data, calculating identifier similarity, and clustering media data according to the identifier similarity include:
[0025] Compare the data identifiers of the media data and calculate the identifier similarity;
[0026] Query the number of storage modules and use the number as the K value;
[0027] Randomly select K centers from the media data, classify other media data based on identification similarity, and synchronously update the centers;
[0028] The process is executed in a loop and the center change is calculated in real time. When the change is less than the preset change threshold, clustering is completed.
[0029] Among them, the change amount uses logo similarity as the evaluation scale.
[0030] As a further solution of the present invention, the step of periodically obtaining device parameters of a receiver within a preset time period for each type of media data and adjusting the data volume of the media data of this type according to the device parameters includes:
[0031] For each type of media data, an adjustment instruction is generated based on a preset frequency timing;
[0032] Each time an adjustment instruction is generated, the device parameters of the recipient within a preset time period are obtained, starting from the current time. The acquisition process includes a prior permission interaction process. The device parameters are used to characterize the performance of the recipient.
[0033] Obtain the mode value range of the device parameter and query the data volume corresponding to the mode value range; the corresponding relationship between the mode value range and the data volume is a preset relationship.
[0034] As a further solution of the present invention, the step of simplifying the media data according to the data volume and sending the simplified media data to the user when a new query request is received includes:
[0035] Determining a retention radius according to the data volume; the retention radius is proportional to the data volume;
[0036] Perform frequency domain conversion on the media data to obtain frequency domain information;
[0037] The frequency domain information within the retention radius is intercepted, and the intercepted frequency domain information is inversely transformed to obtain simplified data;
[0038] When a new query request is received, simplified data of the media data pointed to by the query request is obtained and sent to the user.
[0039] The technical solution of the present invention also provides a media data management system based on big data identification, the system comprising:
[0040] A data identifier extraction module is used to obtain media data, identify the media data, and extract the data identifier; wherein the data identifier extraction process is a random extraction process;
[0041] A data clustering module, configured to compare data identifiers of media data, calculate identifier similarity, and cluster the media data based on the identifier similarity;
[0042] a data volume determination module, configured to periodically obtain device parameters of a receiver within a preset time period for each type of media data, and adjust the data volume of the media data of that type according to the device parameters;
[0043] The media data reduction module is configured to simplify the media data according to the data volume, and send the simplified media data to the user when a new query request is received.
[0044] As a further solution of the present invention: the data identification extraction module includes:
[0045] A data conversion unit, configured to sequentially read media data from a database and convert the media data to obtain a data sequence;
[0046] A data comparison unit is used to compare adjacent data in a data sequence and calculate data similarity;
[0047] A data elimination unit is used to randomly eliminate a piece of data when the data similarity of adjacent data is greater than a preset similarity threshold;
[0048] a loop execution unit, configured to loop execution until the similarities of all adjacent data in the data sequence are less than a preset similarity threshold;
[0049] The extraction execution unit is used to identify each data in the data sequence and extract the data identifier.
[0050] As a further solution of the present invention: the extraction execution unit includes:
[0051] The random selection subunit is used to randomly select data from the data sequence as data to be identified;
[0052] The keyword extraction subunit is used to input the data to be identified into a preset recognition model and extract keywords;
[0053] The statistical subunit is used to count the extracted keywords as data identifiers and synchronously update the selection probability of each data in the data sequence;
[0054] The inner loop sub-unit is used for loop execution until a preset loop exit condition is met; the loop exit condition is a keyword condition;
[0055] The process of determining the selection probability is as follows:
[0056]
[0057] Where, P i is the probability of selecting the i-th data, T i is the characteristic value of the i-th data, N is the total number of data, M is the number of selected data, and σ is the preset parameter; Z j is the calibration value of the j-th selected data, and the calibration value is a preset value.
[0058] As a further solution of the present invention: the data clustering module includes:
[0059] an identification comparison unit, for comparing data identifications of media data and calculating identification similarity;
[0060] A quantity query unit is used to query the number of storage modules and use the number as the K value;
[0061] The classification unit is used to randomly select K centers in the media data, classify other media data according to the similarity of the identification, and synchronously update the centers;
[0062] A calculation unit is used to execute in a loop and calculate the change of the center in real time. When the change is less than a preset change threshold, clustering is completed;
[0063] Among them, the change amount uses logo similarity as the evaluation scale.
[0064] Compared with the existing technology, the beneficial effects of the present invention are: the present invention clusters media data according to its content to obtain multiple categories of media data, queries the average device performance of its audience for each category of media data, simplifies the media data according to the average device performance, and then sends the simplified data to the user, thereby reducing the amount of invalid network resources according to actual conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention.
[0066] Figure 1 The flowchart of the media data management method based on big data identification is shown in FIG.
[0067] Figure 2 This is a block diagram of the first sub-process of the media data management method based on big data identification.
[0068] Figure 3 This is a block diagram of the second sub-process of the media data management method based on big data identification.
[0069] Figure 4 This is a block diagram of the third sub-process of the media data management method based on big data identification.
[0070] Figure 5 This is a fourth sub-process flowchart of the media data management method based on big data identification.
[0071] Figure 6 This is a structural block diagram of a media data management system based on big data identification. DETAILED DESCRIPTION
[0072] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0073] Figure 1 This is a flowchart of a media data management method based on big data identification. In an embodiment of the present invention, a media data management method based on big data identification includes:
[0074] Step S100: Acquire media data, identify the media data, and extract data identifiers; wherein the data identifier extraction process is a random extraction process;
[0075] The media data in this application generally refers to video, which is a set of audio and images. Compared with the image set, the audio data volume is very small and can be stored directly. The focus of this application is the processing process of the image set. After obtaining the media data, the media data is identified and the data identifier can be extracted. When the media data is an image set, the data identifier is the label of each image. To be precise, the label should be a content label, which is used to reflect the content of the image. It should be noted that the data identifier extraction process is a random extraction process. When the media data is an image set, some images are randomly selected from the image set to extract the data identifier. The data identifier can be several keywords.
[0076] Step S200: comparing data identifiers of media data, calculating identifier similarity, and clustering the media data according to the identifier similarity;
[0077] After the media data has gone through the data identification extraction process, the extracted data identification (multiple keywords) can be used as the characteristics of the media data. By comparing the data identifications of any two media data, the similarity between the media data can be calculated, that is, the identification similarity in the above content; using the identification similarity as a reference, the clustering algorithm is applied to the media data to cluster the media data.
[0078] Step S300: For each type of media data, regularly obtain the device parameters of the receiver within a preset time period, and adjust the data volume of the media data of this type according to the device parameters;
[0079] For each type of media data, since the data identifier reflects the image content and the media data is clustered based on the data identifier, the content of the same type of media data is similar, and their audiences (receivers) are roughly similar; for each type of media data, the device parameters of the receiver are counted once every period of time. The device parameters are used to characterize the device performance of the receiver, such as CPU parameters, GPU parameters, memory and running memory, etc., and the data volume of the media data is adjusted based on the device parameters; wherein the adjustment process does not affect the original data, it is based on the copy data of the original data, and the degree of simplification of the copied data is determined according to the data volume. The higher the device performance represented by the device parameters, the larger the data volume and the smaller the degree of simplification. The original data has the largest data volume and is the unsimplified data.
[0080] Step S400: Simplifying the media data according to the data volume, and sending the simplified media data to the user when a new query request is received;
[0081] The media data is simplified according to the data volume, so that each media data in each type of media data corresponds to a simplified copy data. When a new query request is received, the simplified data of the media data pointed to by the query request is obtained and sent to the user.
[0082] The working principle of the above content is that for a batch of media data, it is first clustered according to its content to obtain multiple categories of media data. For each category of media data, the device parameters of its audience are queried to indicate how high the average device performance of the audience is. The media data is simplified according to the average device performance, and the simplified data is sent to the user. In this application, all media data are simplified (relative to the original data). The difference lies in the degree of simplification. If the device performance of the audience of a certain type of media data is low, the degree of simplification is high. The advantage is that the transmission process is simplified without affecting the media effect. Not affecting the media effect means that the corresponding device cannot display the original data at all. For example, the original data is 4K video and the audience's device is an ordinary mobile phone or phone watch. Then transmitting the original data is a waste of resources because the audience's device cannot display it. After simplification, the resource utilization rate becomes higher.
[0083] In addition, the clustering process and the statistical process of device parameters in this application are updated regularly. It is not static, but a dynamic matching solution with high timeliness and high consistency with actual conditions.
[0084] Figure 2 This is a block diagram of the first sub-process of the media data management method based on big data identification. The steps of acquiring media data, identifying the media data, and extracting data identification include:
[0085] Step S101: sequentially reading media data from a database, converting the media data to obtain a data sequence;
[0086] Step S102: Compare adjacent data in the data sequence and calculate data similarity;
[0087] Step S103: When the data similarity of adjacent data is greater than a preset similarity threshold, one data is randomly removed;
[0088] Step S104: cyclically execute until the similarity of all adjacent data in the data sequence is less than a preset similarity threshold;
[0089] Step S105: Identify each data in the data sequence and extract the data identifier.
[0090] In one example of the technical solution of the present invention, media data is read sequentially from a database, and the media data is converted to obtain a data sequence. When the media data is a video ignoring audio, the converted data sequence is an image sequence. For the data in the data sequence, adjacent data are compared sequentially, and data similarity is calculated. When the data similarity of adjacent data is large enough, only one data is retained. The above process is continuously executed in a loop until all adjacent data in the data sequence are dissimilar. Finally, each data in the data sequence is identified and the data identifier is extracted.
[0091] The above process is actually a preprocessing process, which first simplifies the media data, selects only the more important data, and extracts the data identifier. When implementing its function, it only needs to extract the content of the more important data.
[0092] Furthermore, the step of identifying each data in the data sequence and extracting the data identifier includes:
[0093] Randomly select data from the data sequence as the data to be identified;
[0094] Input the data to be identified into the preset recognition model to extract keywords;
[0095] Count the extracted keywords as data identifiers and synchronously update the selection probability of each data in the data sequence;
[0096] The loop is executed until a preset loop exit condition is met; the loop exit condition is a keyword condition.
[0097] The above content is a specific limitation of the random extraction process. Steps S102 to S104 have actually performed a random sampling process, but its randomness is not enough. On this basis, data is randomly selected from the data sequence as the data to be identified, and the data to be identified is input into the preset recognition model to extract keywords, and the extracted keywords are counted as data identifiers.
[0098] Among them, the process of randomly selecting data in the data sequence involves random probability. Each selection and recognition process will update the selection probability once, and each time a data is selected, it will be randomly selected again; the loop is executed, and keywords are continuously extracted until the preset loop exit condition is met; the loop exit condition is the keyword condition, and one of the keyword conditions is that the number of data without new keywords reaches a preset threshold, and the data number is reset to zero when a new keyword appears; its practical meaning is that during the loop execution process, there is a preset threshold number of data that cannot generate new keywords.
[0099] It is worth mentioning that many existing content extraction algorithms can implement the keyword extraction process. The simplest and most effective solution is to directly apply AI, as existing AI has extremely strong comprehension capabilities.
[0100] In addition, the process of determining the selection probability is:
[0101]
[0102] Where, P i is the probability of selecting the i-th data, T i is the characteristic value of the i-th data, N is the total number of data, M is the number of selected data, and σ is the preset parameter; Z j is the calibration value of the j-th selected data, and the calibration value is a preset value.
[0103] The process of determining the selection probability is as follows: for any unselected data, query the influence of other selected data at that location, then sum them up to get the total influence, use the inverse of the total influence as the eigenvalue of that location, calculate the ratio of the eigenvalue of each location to the sum of the eigenvalues of all unselected data, and then get the selection probability; the meaning of influence is that the greater the influence, the smaller the eigenvalue, and the smaller the probability of the data at that location being selected, which means that if the adjacent location of a location is selected, then its selection probability will be very small.
[0104] Specifically, regarding the impact, The term is a weight term, which means that the smaller the span between two points (the jth selected data and the ith data), the smaller the impact. jThe value of is a preset value. The simplest way is to set it directly to 1, which means that the calibration value of all data is the same. However, in actual situations, it can also be linked to the extracted identifier. For example, query the extracted identifier corresponding to the j-th selected data, and query the total number of extracted identifiers. The larger the total number, the larger the calibration value and the greater the impact. The effect of this process is that if a identifier is common, the probability of selecting the data around it will be very small, and it will almost not be selected. If a identifier is uncommon, the probability of selecting the data around it is still very high, and it will still be selected.
[0105] Figure 3 This is a second sub-flow diagram of the media data management method based on big data identification. The steps of comparing the data identifications of the media data, calculating the identification similarity, and clustering the media data according to the identification similarity include:
[0106] Step S201: Compare data identifiers of media data and calculate identifier similarity;
[0107] Step S202: Query the number of storage modules and use the number as the K value;
[0108] Step S203: randomly select K centers from the media data, classify other media data according to identification similarity, and synchronously update the centers;
[0109] Step S204: cyclically execute and calculate the change of the center in real time. When the change is less than a preset change threshold, clustering is completed.
[0110] Among them, the change amount uses logo similarity as the evaluation scale.
[0111] In an example of the technical solution of the present invention, the clustering process is limited to the K-means algorithm, the data identifiers of the media data are compared, and the identifier similarity is calculated. Since the data identifier is a keyword, the identifier similarity can be expressed as the ratio of the number of intersections to the number of unions; then, the number of storage modules is queried, which indicates how many categories all the media data should be divided into. This is related to the storage architecture. The more storage areas, the larger the number; finally, K centers are randomly selected from the media data, and other media data are classified according to the identifier similarity. The centers are updated synchronously, and the execution is cyclic and the change in the centers is calculated in real time. When the change is less than a preset change threshold, clustering is completed; wherein, the change uses identifier similarity as an evaluation scale. A simpler way is that the identifier similarity of the centers determined twice adjacently is less than a preset threshold.
[0112] Figure 4This is a block diagram of a third sub-process of a media data management method based on big data identification. For each type of media data, the steps of periodically obtaining the device parameters of the receiver within a preset time period and adjusting the data volume of the media data of that type according to the device parameters include:
[0113] Step S301: Generate an adjustment instruction for each type of media data based on a preset frequency timing;
[0114] Step S302: Each time an adjustment instruction is generated, the device parameters of the recipient within a preset period are obtained, starting from the current time. The acquisition process includes a prior permission interaction process. The device parameters are used to characterize the performance of the recipient.
[0115] Step S303: Obtain the mode value range of the device parameter, and query the data volume corresponding to the mode value range; the corresponding relationship between the mode value range and the data volume is a preset relationship.
[0116] In an example of the technical solution of the present invention, for each type of media data, an adjustment instruction is generated based on a preset frequency. Each time an adjustment instruction is generated, the device parameters of the receiver within a preset time period are obtained starting from the current moment. Since the device parameters are the data of the receiver, it is necessary to obtain permission first and then obtain them. The obtained device parameters can directly use the evaluation score of existing software. Many existing software will evaluate the performance of the device, commonly known as "running score". The larger the evaluation score, the higher the device performance.
[0117] Finally, obtain the mode range of the device parameters, and query the data volume corresponding to the mode range. The correspondence between the mode range and the data volume is a preset relationship. A mapping table can be pre-determined by the administrator of the subject executing this method. For the subject executing this method, it is the default known data. The concept of the mode range is an extension of the mode. When creating the mapping table, the relationship between the mode range and the data volume is determined. For example, 0 to 10,000 points corresponds to how much data volume, 10,000-50,000 points corresponds to how much data volume, and so on. In actual use, query which mode range the device parameter belongs to and directly read the data volume.
[0118] Figure 5 This is a fourth sub-flow diagram of the media data management method based on big data identification, wherein the steps of simplifying the media data according to the data volume and sending the simplified media data to the user when a new query request is received include:
[0119] Step S401: determining a retention radius according to the data volume; the retention radius is proportional to the data volume;
[0120] Step S402: Perform frequency domain conversion on the media data to obtain frequency domain information;
[0121] Step S403: intercepting the frequency domain information within the retention radius, and performing inverse conversion on the intercepted frequency domain information to obtain simplified data;
[0122] Step S404: When a new query request is received, simplified data of the media data pointed to by the query request is obtained and sent to the user.
[0123] In an example of the technical solution of the present invention, a data simplification process is described. The retention radius is determined according to the data volume, the media data is converted into the frequency domain to obtain frequency domain information, the frequency domain information within the retention radius is intercepted with the origin as the center of the circle, and the intercepted frequency domain information is inversely converted to obtain simplified data. From the above content, it can be seen that the larger the data volume, the larger the retention radius, the more data retained, and the smaller the degree of simplification; finally, when a new query request is received, the simplified data of the media data pointed to by the query request is obtained and sent to the user.
[0124] Figure 6 : is a structural block diagram of a media data management system based on big data identification. In an embodiment of the present invention, a media data management system based on big data identification, the system 10 includes:
[0125] The data identification extraction module 11 is used to obtain media data, identify the media data, and extract the data identification; wherein the data identification extraction process is a random extraction process;
[0126] A data clustering module 12 is configured to compare data identifiers of media data, calculate identifier similarity, and cluster the media data based on the identifier similarity;
[0127] The data volume determination module 13 is configured to periodically obtain device parameters of a receiver within a preset time period for each type of media data, and adjust the data volume of the media data of this type according to the device parameters;
[0128] The media data reduction module 14 is configured to simplify the media data according to the data volume, and send the simplified media data to the user when a new query request is received.
[0129] Furthermore, the data identification extraction module 11 includes:
[0130] A data conversion unit, configured to sequentially read media data from a database and convert the media data to obtain a data sequence;
[0131] A data comparison unit is used to compare adjacent data in a data sequence and calculate data similarity;
[0132] A data elimination unit is used to randomly eliminate a piece of data when the data similarity of adjacent data is greater than a preset similarity threshold;
[0133] a loop execution unit, configured to loop execution until the similarities of all adjacent data in the data sequence are less than a preset similarity threshold;
[0134] The extraction execution unit is used to identify each data in the data sequence and extract the data identifier.
[0135] Specifically, the extraction execution unit includes:
[0136] The random selection subunit is used to randomly select data from the data sequence as data to be identified;
[0137] The keyword extraction subunit is used to input the data to be identified into a preset recognition model and extract keywords;
[0138] The statistical subunit is used to count the extracted keywords as data identifiers and synchronously update the selection probability of each data in the data sequence;
[0139] The inner loop sub-unit is used for loop execution until a preset loop exit condition is met; the loop exit condition is a keyword condition;
[0140] The process of determining the selection probability is as follows:
[0141]
[0142] Where, P i is the probability of selecting the i-th data, T i is the characteristic value of the i-th data, N is the total number of data, M is the number of selected data, and σ is the preset parameter; Z j is the calibration value of the j-th selected data, and the calibration value is a preset value.
[0143] Furthermore, the data clustering module 12 includes:
[0144] an identification comparison unit, for comparing data identifications of media data and calculating identification similarity;
[0145] A quantity query unit is used to query the number of storage modules and use the number as the K value;
[0146] The classification unit is used to randomly select K centers in the media data, classify other media data according to the similarity of the identification, and synchronously update the centers;
[0147] A calculation unit is used to execute in a loop and calculate the change of the center in real time. When the change is less than a preset change threshold, clustering is completed;
[0148] Among them, the change amount uses logo similarity as the evaluation scale.
[0149] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A media data management method based on big data identification, characterized in that: The method comprises: Acquire media data, identify the media data, and extract a data identifier; wherein the data identifier is extracted in a random manner; Comparing data identifiers of the media data, calculating identifier similarity, and clustering the media data based on the identifier similarity; For each type of media data, regularly obtain the device parameters of the receiver within a preset time period, and adjust the data volume of the media data of this type according to the device parameters; The media data is simplified according to the data volume, and when a new query request is received, the simplified media data is sent to the user.
2. The media data management method based on big data identification according to claim 1, characterized in that: The steps of acquiring media data, identifying the media data, and extracting a data identifier include: Reading media data in sequence from a database, converting the media data to obtain a data sequence; Compare adjacent data in the data sequence and calculate data similarity; When the data similarity of adjacent data is greater than the preset similarity threshold, one data is randomly removed; The loop is executed until the similarity of all adjacent data in the data sequence is less than the preset similarity threshold; Identify each data in the data sequence and extract the data identifier.
3. The media data management method based on big data identification according to claim 2, characterized in that: The step of identifying each data in the data sequence and extracting the data identifier includes: Randomly select data from the data sequence as the data to be identified; Input the data to be identified into the preset recognition model to extract keywords; Count the extracted keywords as data identifiers and synchronously update the selection probability of each data in the data sequence; The loop is executed until a preset loop exit condition is met; the loop exit condition is a keyword condition; The process of determining the selection probability is as follows: Where, P i is the probability of selecting the i-th data, T i is the characteristic value of the i-th data, N is the total number of data, M is the number of selected data, and σ is the preset parameter; j is the calibration value of the j-th selected data, and the calibration value is a preset value.
4. The media data management method based on big data identification according to claim 1, characterized in that: The steps of comparing data identifiers of media data, calculating identifier similarity, and clustering media data according to the identifier similarity include: Compare the data identifiers of the media data and calculate the identifier similarity; Query the number of storage modules and use the number as the K value; Randomly select K centers from the media data, classify other media data based on identification similarity, and synchronously update the centers; The process is executed in a loop and the center change is calculated in real time. When the change is less than the preset change threshold, clustering is completed. Among them, the change amount uses logo similarity as the evaluation scale.
5. The media data management method based on big data identification according to claim 1, characterized in that: The step of regularly obtaining device parameters of a receiver within a preset time period for each type of media data and adjusting the data volume of the media data of this type according to the device parameters includes: For each type of media data, an adjustment instruction is generated based on a preset frequency timing; Each time an adjustment instruction is generated, the device parameters of the recipient within a preset time period are obtained, starting from the current time. The acquisition process includes a prior permission interaction process. The device parameters are used to characterize the performance of the recipient. Obtain the mode value range of the device parameter and query the data volume corresponding to the mode value range; the corresponding relationship between the mode value range and the data volume is a preset relationship.
6. The media data management method based on big data identification according to claim 1, characterized in that: The step of simplifying the media data according to the data volume and sending the simplified media data to the user when a new query request is received includes: Determining a retention radius according to the data volume; the retention radius is proportional to the data volume; Perform frequency domain conversion on the media data to obtain frequency domain information; The frequency domain information within the retention radius is intercepted, and the intercepted frequency domain information is inversely transformed to obtain simplified data; When a new query request is received, simplified data of the media data pointed to by the query request is obtained and sent to the user.
7. A media data management system based on big data identification, characterized in that: The system comprises: A data identifier extraction module is used to obtain media data, identify the media data, and extract the data identifier; wherein the data identifier extraction process is a random extraction process; A data clustering module, configured to compare data identifiers of media data, calculate identifier similarity, and cluster the media data based on the identifier similarity; a data volume determination module, configured to periodically obtain device parameters of a receiver within a preset time period for each type of media data, and adjust the data volume of the media data of that type according to the device parameters; The media data reduction module is configured to simplify the media data according to the data volume, and send the simplified media data to the user when a new query request is received.
8. The media data management system based on big data identification according to claim 7, characterized in that: The data identification extraction module includes: A data conversion unit, configured to sequentially read media data from a database and convert the media data to obtain a data sequence; A data comparison unit is used to compare adjacent data in a data sequence and calculate data similarity; A data elimination unit, configured to randomly eliminate a piece of data when the data similarity of adjacent data is greater than a preset similarity threshold; a loop execution unit, configured to loop execution until the similarities of all adjacent data in the data sequence are less than a preset similarity threshold; The extraction execution unit is used to identify each data in the data sequence and extract the data identifier.
9. The media data management system based on big data identification according to claim 8, characterized in that: The extraction execution unit includes: A random selection subunit is used to randomly select data from a data sequence as data to be identified; The keyword extraction subunit is used to input the data to be identified into a preset recognition model and extract keywords; The statistical subunit is used to count the extracted keywords as data identifiers and synchronously update the selection probability of each data in the data sequence; The inner loop sub-unit is used for loop execution until a preset loop exit condition is met; the loop exit condition is a keyword condition; The process of determining the selection probability is as follows: Where, P i is the probability of selecting the i-th data, T i is the characteristic value of the i-th data, N is the total number of data, M is the number of selected data, and σ is the preset parameter; j is the calibration value of the j-th selected data, and the calibration value is a preset value.
10. The media data management system based on big data identification according to claim 7, characterized in that: The data clustering module includes: an identification comparison unit, for comparing data identifications of media data and calculating identification similarity; A quantity query unit is used to query the number of storage modules and use the number as the K value; The classification unit is used to randomly select K centers in the media data, classify other media data according to the similarity of the identification, and synchronously update the centers; A calculation unit is used to execute in a loop and calculate the change of the center in real time. When the change is less than a preset change threshold, clustering is completed; Among them, the change amount uses logo similarity as the evaluation scale.