A method and system for detecting and processing genuine software based on big data analysis
Through big data analysis and deep learning models, combined with multimodal neural networks and self-attention mechanisms, the problems of cumbersome manual operations and poor flexibility in existing genuine software detection methods are solved, and efficient and accurate pirated software identification and management are achieved.
Patent Information
- Application Number
- CN202411350700.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2044-09-26
AI Technical Summary
Existing methods for detecting genuine software rely on cumbersome manual operations, have poor detection flexibility, and are inefficient. They are unable to cope with the management of massive terminal devices in large enterprises or government agencies, and lack in-depth analysis of software usage behavior.
It adopts big data analysis and deep learning technology, obtains multi-dimensional risk features from software installation source information and online service logs, uses deep learning models to automatically identify piracy risks, and combines multimodal neural networks and self-attention mechanisms for feature fusion to achieve a comprehensive understanding and identification of software behavior.
It achieves efficient and accurate genuine software detection, reduces manual intervention, improves management efficiency, adapts to complex software environments, dynamically responds to new piracy methods, and improves detection accuracy and robustness.
Smart Images

Figure CN119203077B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of software management technology, and in particular to a method and system for detecting and processing genuine software based on big data analysis. Background Art
[0002] To address the issue of software piracy, governments and software companies around the world are actively taking measures to promote the use of genuine software. Currently, in the field of software legalization management, a number of tools based on manual review and simple scanning have been put on the market to help organizations and enterprises manage software installation and usage.
[0003] Currently, the detection and management of genuine software is primarily accomplished through the following steps: first, software installation information on terminal devices is collected through manual recording or automated scanning; then, manual analysis is performed to confirm the software's legality and authorization status; and finally, a report is generated for further compliance checks by enterprise or institutional managers. Therefore, manual work accounts for a significant portion of the process; whether it is entering and classifying software information or verifying authorization status, it relies on a significant amount of manual effort, which not only increases management costs but also significantly reduces inspection efficiency. However, in large enterprises or government agencies, with thousands of computer terminals requiring individual inspection, manual inspection and management methods are often insufficient and prone to data omissions and errors.
[0004] In addition, the functions of current genuine software checking tools are relatively simple, mainly focusing on simple scanning and data comparison, relying on hard-coded rules to determine authorization status, lacking in-depth analysis and understanding of software usage behavior, and seriously lacking flexibility and intelligence, resulting in detection results that are often unsatisfactory.
[0005] To address the above issues, the industry has not yet proposed a better technical solution. Summary of the Invention
[0006] The present application provides a method, system, storage medium, computer program product and electronic device for detecting and processing genuine software based on big data analysis, which is used to at least solve the problems of cumbersome manual operation, poor detection flexibility and low efficiency in the current related technologies, and can efficiently, accurately and intelligently perform genuine software detection and management.
[0007] In the first aspect, an embodiment of the present application provides a method for detecting and processing genuine software based on big data analysis, including: obtaining the installation source information of each installed software in a client device to be detected; verifying the installation source information of each installed software based on a software installation authorization library to identify whether each installed software has a potential piracy risk; the software authorization library contains multiple authorized genuine software names and corresponding software authorized installation channels; for each risky installed software identified as having a potential piracy risk, obtaining the online service log of the risky installed software, and extracting multidimensional risk features from the online service log; the multidimensional risk features include software online update response frequency, software functional module usage frequency, software plug-in loading information and software user behavior information; inputting the multidimensional risk features of each risky installed software into a piracy identification model to determine whether there is a piracy risk accordingly; the piracy identification model adopts a deep learning model.
[0008] In the second aspect, an embodiment of the present application provides a genuine software detection and processing system based on big data analysis, including: an installation source acquisition unit, used to obtain the installation source information of each installation software in the client device to be detected; a piracy risk initial identification unit, used to verify the installation source information of each installation software based on the software installation authorization library, so as to identify whether each installation software has a potential piracy risk; the software authorization library contains multiple authorized genuine software names and corresponding software authorization installation channels; a service log acquisition unit, used to obtain the online service log of the risky installation software for each risky installation software identified as having a potential piracy risk, and extract multi-dimensional risk features from the online service log; the multi-dimensional risk features include the software online update response frequency, the software function module usage frequency, the software plug-in loading information and the software user behavior information; a piracy risk determination unit, used to input the multi-dimensional risk features of each risky installation software into a piracy identification model, so as to determine whether there is a piracy risk accordingly; the piracy identification model adopts a deep learning model.
[0009] According to a third aspect, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the steps of the genuine software detection and processing method based on big data analysis of any embodiment of the present application.
[0010] In a fourth aspect, an embodiment of the present application provides a storage medium on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the genuine software detection and processing method based on big data analysis of any embodiment of the present application are implemented.
[0011] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the genuine software detection and processing method based on big data analysis of any embodiment of the present application.
[0012] The method and system for detecting and processing genuine software based on big data analysis provided by this application can produce at least the following technical effects:
[0013] (1) Through big data analysis and deep learning technology, traditional manual inspection and simple rule judgment are transformed into automated multi-dimensional data analysis. The software installation source information is initially verified using the software installation authorization library, and the piracy risk of potential risk software is further automatically identified based on deep learning. This greatly reduces the reliance on manual review and is particularly suitable for the management of massive terminal devices in large enterprises and institutions, thereby effectively improving management efficiency and avoiding missed detection and misjudgment problems in manual operations.
[0014] (2) In this technical solution, instead of relying solely on traditional simple scanning and hard-coded rules to determine the software's authorization status, a piracy identification model based on deep learning is introduced. By extracting multi-dimensional risk characteristics of the software (such as online update response frequency, functional module usage frequency, plug-in loading information, etc.) through online service logs, an in-depth analysis of software usage behavior can be conducted. This allows for flexible response to various complex software usage scenarios, significantly improving the intelligence and flexibility of detection and avoiding the limitations of single rule detection.
[0015] (3) Deep learning models can capture complex patterns and associations in large-scale data and effectively identify potential risks that are difficult to detect with rule-based judgments. By learning the correlations and regularities of various risk characteristics, the model can autonomously optimize its recognition capabilities when faced with unknown software environments or ever-changing software usage behaviors, thereby improving the overall detection effect. Therefore, through this technical solution, it is possible to achieve stronger dynamic adaptability, timely identify new piracy methods and technologies, and thus effectively respond to the ever-changing pirated software environment and maintain a high recognition accuracy and robustness.
[0016] Through this technical solution, through the application of big data analysis and deep learning models, massive data can be quickly analyzed and processed, realizing automatic detection of the authenticity of installed software, greatly reducing the involvement of manual operations, and helping enterprises or institutions to significantly reduce labor costs. In addition, this detection system framework has good scalability and can adapt to the complex software environment requirements of large enterprises or government agencies. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 A flowchart illustrating an example of a method for detecting and processing genuine software based on big data analysis according to an embodiment of the present application is shown;
[0019] Figure 2 A schematic diagram showing the structural connection of an example of a piracy identification model according to an embodiment of the present application is shown;
[0020] Figure 3 An operational flow chart of an example of extracting a time series feature matrix according to an embodiment of the present application is shown;
[0021] Figure 4 An operational flowchart of an example of adaptively determining convolution kernel scale weights based on information entropy weight calculation according to an embodiment of the present application is shown;
[0022] Figure 5 A structural block diagram of an example of a genuine software detection and processing system based on big data analysis according to an embodiment of the present application is shown;
[0023] Figure 6 This is a schematic structural diagram of an embodiment of an electronic device of the present application. DETAILED DESCRIPTION
[0024] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0025] In the technical solutions of this application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved shall comply with the provisions of relevant laws and regulations and shall not violate public order and good morals.
[0026] Figure 1 A flowchart of an example of a method for detecting and processing genuine software based on big data analysis according to an embodiment of the present application is shown.
[0027] Regarding the executor of the method of the embodiment of the present application, it can be any controller or processor with computing or processing capabilities. Through big data analysis, multi-dimensional risk feature extraction and the introduction of deep learning models, a set of intelligent and efficient genuine software detection and processing methods are formed, which can effectively improve detection accuracy, reduce manual intervention and reduce management costs.
[0028] In some examples, it can be integrated into an electronic device or terminal through software, hardware, or a combination of software and hardware, and the type of terminal or electronic device can be diverse, such as a mobile phone, tablet computer, or desktop computer, etc.
[0029] like Figure 1 As shown, in step S110, the installation source information of each installed software in the client device to be detected is obtained.
[0030] In some embodiments, the detection system (or detection tool) can first obtain relevant information about the installed software on the client device, especially the installation source information of the software, in an automated manner using a scanning tool (such as API call, system log analysis or registry parsing). The types of installation source information can be diverse, such as download channels (such as official website download, third-party site download, USB installation, network sharing, etc.), installation time, installation path, etc. Here, it is necessary to accurately sample and scan the source information of each installed software in the client device to ensure that all installed software can be fully covered, which can avoid mistakes or omissions in manual recording and improve the efficiency and accuracy of detection.
[0031] In step S120, the installation source information of each installation software is verified based on the software installation authorization library to identify whether each installation software has a potential piracy risk.
[0032] It should be noted that the Software Authorization Library contains multiple authorized genuine software titles and corresponding authorized software installation channels. The Software Authorization Library pre-stores a list of legal software, including genuine software titles and their corresponding legal installation channels. In addition, system administrators can regularly update and maintain the Software Authorization Library to ensure that the genuine software information in it is up-to-date and covers a wide range of authorized channels (such as official websites and trusted app stores).
[0033] Here, the installation source information for each software item is compared using data from the software authorization library. The authorization library contains the names of legally licensed software and their corresponding installation channels. By comparing the software name and installation channel, we determine whether the software is genuinely licensed. For example, if a piece of software is installed through an unauthorized third-party channel, it will be marked as potentially piracy-prone. This comparison with the authorization library allows for rapid screening of software with high accuracy and reduced false positives. It also avoids the need for subsequent piracy risk analysis for all installed software, improving piracy detection efficiency while reducing resource consumption.
[0034] In step S130 , for each risky installation software identified as having a potential piracy risk, an online service log of the risky installation software is obtained, and a multi-dimensional risk feature is extracted from the online service log.
[0035] Here, for potentially pirated software identified as having non-compliant installation sources, online service logs are obtained through network interfaces. These logs include communication records between the software and the server, update response logs, and plugin loading logs during use. Multidimensional risk signatures are then extracted from these logs. These signatures include the frequency of online update responses, the frequency of software functional module usage, plugin loading information, and user behavior.
[0036] The description of the online update response frequency can be used to determine whether the software is frequently updated and whether the updates are consistent with the update cycle of genuine software. Specifically, genuine software usually has regular update push notifications, and users will upgrade according to the update frequency announced by the manufacturer. Pirated software may lack this update service or have abnormal update response behavior (such as irregularly skipping updates or failing to update). Therefore, analyzing users' update response behavior can help identify software piracy risks.
[0037] The frequency of functional module usage can be analyzed to determine whether the frequency of use of the software's core functions matches the normal usage of legitimate software. Specifically, the functional modules of legitimate software are generally stable and comprehensive, allowing users to access and use all functional modules. Pirated software, on the other hand, may lack some functionality or operate unstably. By recording data such as the number of calls to software functional modules and the success rate of their execution, it is possible to detect whether users are chronically lacking access to advanced features or whether they frequently experience abnormal behaviors such as crashes and errors.
[0038] The description of plugin loading information can monitor whether the software has loaded unauthorized plugins or extension modules. Specifically, some pirated software may bypass the security mechanisms of legitimate software through plug-ins, illegal patches, or cracking tools. These tools often tamper with the software's file structure or load illegal plugins, causing the software to behave differently from legitimate software during startup or operation. By analyzing the file structure, memory usage, and plugin loading status, it is possible to extract characteristics of illegal plugin use.
[0039] Regarding software user behavior information, it can record user operations and be used to assess whether there are abnormal operations or unusual usage patterns. Specifically, legitimate users have relatively fixed operating habits, such as frequency of use and operating time periods. Pirated users may have more unusual behavior, such as frequent device switching, cross-region use, and abnormal operating time periods. By modeling the time and frequency of user operation behaviors, it helps to identify risks of pirated software that do not conform to the behavior patterns of legitimate users.
[0040] Therefore, the introduction of multi-dimensional risk features designed with the above-mentioned feature dimensions can better understand the software's operating mode and user behavior, further improve the accuracy of piracy identification, and help identify software that appears to be genuine but may actually contain pirated features.
[0041] In step S140, the multi-dimensional risk features of each risky installed software are respectively input into a piracy identification model to determine whether there is a piracy risk. The piracy identification model adopts a deep learning model.
[0042] Here, the piracy identification model can be trained in advance using a data sample set containing a large amount of historical data and annotation information, so that it can accurately distinguish the behavioral patterns of genuine software and pirated software.
[0043] In some implementations, during the inference phase, the piracy identification model can analyze the input multi-dimensional risk features to generate a piracy risk score or probability value. If the probability value exceeds a preset threshold, the software is deemed pirated. Furthermore, deep learning models can continuously optimize themselves, updating with new data to enhance their accuracy and generalization capabilities.
[0044] Through the embodiments of the present application, the entire process from data collection to risk identification is highly automated, significantly improving detection efficiency. First, the installation source information is used for global screening and verification to identify risky installed software on the client. Then, a deep learning model is combined with multi-dimensional feature analysis to conduct in-depth analysis of each risky installed software to identify potential pirated software, which can effectively improve detection accuracy. In addition, by continuously training the deep learning model and updating the software authorization library, the detection system can quickly adapt to new software and piracy methods that appear on the market, while supporting the expansion needs of large-scale enterprise environments.
[0045] In some examples of the embodiments of the present application, the piracy identification model is a multimodal neural network that comprehensively utilizes multiple types of risk features, specifically performing comprehensive analysis from two aspects: the static features of the temporal relationship and the nonlinear relationship, to predict the final risk probability value of pirated software.
[0046] Figure 2 A schematic diagram of the structural connection of an example of a piracy identification model according to an embodiment of the present application is shown.
[0047] like Figure 2 As shown, the multimodal neural network 200 includes a CNN (Convolutional Neural Network) module 210, an MLP (Multi-Layer Perceptron) module 220, a feature fusion module 230 and a classification module 240.
[0048] The CNN module 210 is used to extract the local dependency between the software online update response frequency and the software functional module usage frequency in the multi-dimensional risk features to generate corresponding temporal pattern features.
[0049] Specifically, the CNN module reconstructs the time series structure of the software's online update response frequency and the frequency of software functional module usage, adapting it to convolution operations. Next, multiple convolutional layers are used to extract local features from the time series data. The convolution kernel uses a sliding window to capture local feature correlations and frequency variation patterns, such as periodic fluctuations in the online update response frequency and peaks and troughs in the frequency of functional module usage. By stacking convolutional layers, higher-level time series pattern features can be extracted layer by layer, ultimately generating a time series feature.
[0050] The MLP module 220 is used to extract the nonlinear relationship between the software plug-in loading information and the software user behavior information in the multi-dimensional risk feature to generate corresponding static pattern features.
[0051] Specifically, the MLP module primarily processes the static information contained in multidimensional risk features, processing plug-in loading information and user behavior information to extract the complex, high-dimensional, nonlinear patterns between these features. Specifically, the MLP module consists of an input layer, a hidden layer, and an output layer. Through the input layer, the MLP module inputs and embeds static features into a high-dimensional space. Through the hidden layer, the MLP module's multi-layer, fully connected structure effectively captures the nonlinear relationships between complex features. For example, plug-in loading information may have complex underlying connections with user operational behavior, and the hidden layer extracts these deep-level connections layer by layer. Ultimately, the MLP module outputs a feature vector representing the combined pattern of learned static features, namely the static pattern feature.
[0052] It should be noted that while software plug-in loading information and user behavior data are static, they often reveal complex software usage habits and behavioral patterns. The MLP module can capture the complex nonlinear relationships between plug-in loading and user behavior, such as the combination of different plug-ins and the order in which they are loaded, enhancing the ability to identify pirated software. This allows for better identification of potential piracy risks arising from complex behavioral patterns, particularly in scenarios such as plug-in loading failures and abnormal user behavior. The model can accurately detect the characteristics of pirated software, reducing false positives and missed negatives.
[0053] By combining CNN and MLP modules, the system can simultaneously process temporal features (such as online update response frequency and functional module usage frequency) and static features (such as plugin loading information and user operation behavior). This ensures that local patterns and global nonlinear relationships between features are fully mined and modeled, avoiding the problem of a single model being unable to capture complex behavioral patterns. As a result, deep mining enables the system to comprehensively analyze software behavior, thereby improving the accuracy of identifying pirated software, especially demonstrating strong recognition capabilities for software with abnormal functions, abnormal update responses, and unstable plugin loading.
[0054] The feature fusion module 230 is used to fuse the temporal pattern features and the static pattern features to generate corresponding fused features.
[0055] Here, the feature fusion module can fuse temporal and static pattern features using various unrestricted methods, such as simple feature concatenation or weighted fusion. This allows the fused features to incorporate comprehensive information from different dimensions and time layers, providing richer contextual feature data. By effectively fusing temporal and static pattern features, we achieve the comprehensive utilization of multimodal information and a comprehensive understanding of software behavior, effectively addressing different types of pirated software and user behavior patterns.
[0056] The classification module 240 is used to classify the fused features through a fully connected layer to output a risk probability value that the risky installed software is pirated software.
[0057] Here, the classification module can employ a typical fully connected network to classify the fused features. Specifically, the high-dimensional fused features from the fully connected network are processed through several layers, and then, in the final layer, a sigmoid activation function is used to generate a risk probability value indicating that the risky installed software is pirated. This risk probability value is then used to determine whether the software is pirated. This deep learning-based classification process is therefore highly efficient when processing large amounts of data, rapidly generating piracy detection results and providing real-time pirated software detection services to businesses and users.
[0058] Through the embodiments of this application, using a multimodal neural network structure, the system can be expanded based on the increase or change of feature dimensions (such as adding new behavioral features or plug-in loading modes). At the same time, through the continuous training mechanism of deep learning, the model can adapt to changes in software behavior in different usage scenarios. As a result, the detection system has good scalability and adaptability, and can be continuously optimized and upgraded as the software market and user behavior change, effectively responding to the detection needs of new pirated software.
[0059] In some examples of the present application, the feature fusion module 230 uses a feature fusion module based on the self-attention mechanism. By introducing the self-attention mechanism, the feature fusion module can perform fine-grained interactive analysis of temporal features and static features, so that feature fusion is no longer a simple linear combination, but rather dynamically captures the complex relationship between features through the attention mechanism.
[0060] Specifically, the temporal pattern features and static pattern features corresponding to the first risky installed software are received, and the feature matrix of the temporal pattern features is: T is the total number of time steps, F1 is the dimension of the time series feature, and the feature matrix of the static pattern feature is F2 is the dimension of static features.
[0061] The temporal pattern features and static pattern features are projected separately through the linear transformation matrix to generate the corresponding query vector, key vector and value vector:
[0062]
[0063] Where, is the linear transformation weight matrix of query, key, and value of time series features; is the linear transformation weight matrix of the query, key, and value of the static feature, d is the dimension of the potential feature, which is used to unify the representation of all input features; Q s ,K s,V s Represent the query vector, key vector and value vector corresponding to the time series features, Q p ,K p ,V p They represent the query vector, key vector, and value vector corresponding to the static features respectively.
[0064] For the interaction between temporal features and static features, attention calculations are performed separately:
[0065] Z s =Attention(Q s ,K s ,V s ), formula (3)
[0066] Z p =Attention(Q p ,K p ,V p ), formula (4)
[0067]
[0068] Where Attention(·) represents the dot product attention calculation mechanism; is a scaling factor to prevent the dot product value from being too large and affecting the gradient calculation; softmax represents the softmax activation function, which converts the correlation value into a weight coefficient through the Softmax activation operation. The weight coefficient is used to perform a weighted summation on V to generate the fused feature representation; K T represents the transpose of the key vector K; Z s is the temporal feature self-attention representation; Z p is the static feature self-attention representation.
[0069] Here, the self-attention mechanism automatically assigns weights to each feature based on its importance in the current scenario. For features that pose a higher risk of piracy, such as unusual patterns in the frequency of online software updates, the model can assign higher weights, ensuring that these key features have a greater impact on the final judgment. Furthermore, the self-attention mechanism's dynamic weight adjustment ensures that the model is more sensitive to key features. For example, if a piece of software has not been properly updated for an extended period or if user operation frequency is unusually high, the model will adaptively focus on these features, thereby improving the accuracy of piracy detection.
[0070] By mixing the query vector and key vector of temporal features and static features, we can further capture the interaction between them:
[0071] Z sp =Attention(Q s ,Kp ,V p ), formula (6)
[0072] Z ps =Attention(Q p ,K s ,V s ), formula (7)
[0073] Where Z sp Represents the first interactive attention feature of temporal features to static features, Z ps Represents the second interactive attention feature of static features to temporal features.
[0074] It should be noted that among multimodal features, the interaction complexity between temporal and static features is often quite high. Through the self-attention mechanism, the model can establish deep connections between temporal and static features, capturing potential interaction patterns between them. For example, fluctuations in online update frequency may be linked to plugin loading failures or frequent user operations. Therefore, the collaborative analysis between temporal and static features using the self-attention mechanism helps the model better understand the logic behind certain complex behavioral patterns. For example, if a plugin fails to load and the update frequency is abnormal, the detection system can capture the deep connections between these behavioral patterns, significantly increasing the likelihood of identifying the software as pirated.
[0075] Fuse the temporal feature self-attention representation, the static feature self-attention representation, the first interactive attention feature and the second interactive attention feature:
[0076] Z final =Z s +Z p +Z sp +Z ps , formula (8)
[0077] Where Z final Indicates the fusion feature corresponding to the first risky installed software.
[0078] The present invention utilizes a self-attention mechanism to give the feature fusion module greater adaptability, automatically adjusting the importance of features as data dynamically changes. This makes the model more robust and scalable across diverse software usage scenarios. Furthermore, by capturing more accurate feature information through dynamic weighting and multi-dimensional interaction, the output of feature fusion has higher information density and relevance, resulting in higher accuracy and lower false positive rates in piracy detection results.
[0079] In some examples of the embodiments of the present application, the CNN module adopts a multi-scale convolutional network module, which introduces multiple convolution kernels of different scales to simultaneously capture short-term fine-grained and long-term coarse-grained feature behavior patterns.
[0080] Specifically, convolution kernels of different scales act on the input feature matrix to generate multiple different feature maps that reflect short-term and long-term behavior patterns. For example, smaller convolution kernels can better capture feature changes in a short period of time, while larger convolution kernels can capture features over a long period of time.
[0081]
[0082] Where, F i is through a size of k i ×k i The feature matrix generated by the convolution kernel is Indicates the scale is k i ×k i The convolution kernel weight matrix, X t Represents the input time series feature matrix, Conv(·) represents the convolution process; N is the total number of convolution kernels, α i is of scale k i ×k i The convolution kernel scale weight corresponding to the convolution kernel.
[0083] Here, the features at different scales are fused by summing operation, and the convolution kernel scale weight α is used i Controlling the influence of each scale in the final fusion feature enables the model to flexibly adjust the importance of each scale, thereby better capturing the behavioral patterns of pirated software.
[0084] It should be noted that the convolution kernel scale weight can be determined in a variety of ways. On the one hand, it can be pre-set through user input parameters. On the other hand, it can be used as a learnable parameter and adjusted during the model training iteration process, and all of them fall within the scope of implementation of the embodiments of this application.
[0085] The embodiments of this application utilize a multi-scale convolutional network module for multi-scale convolution and feature fusion, enabling the capture of multi-layered behavioral patterns ranging from short-term anomalies to long-term stagnation. This allows for the simultaneous identification of behavioral features at different time scales, improving the model's ability to detect complex behavioral patterns and, in particular, increasing sensitivity to subtle anomalies exhibited by pirated software at different times. Furthermore, through the use of convolution kernel scale weights, the model can automatically adjust the importance of different convolution kernel scales based on actual scenarios, enhancing the adaptability of feature representation.
[0086] Figure 3An operational flowchart of an example of extracting a time series feature matrix according to an embodiment of the present application is shown.
[0087] In step S310 , time series analysis is performed on the online service log according to a preset time window size to extract a time series feature group.
[0088] Here, the timing feature group contains timing features corresponding to a preset number of time windows. The timing features include the software online update response frequency and the usage frequency of software functional modules. Therefore, feature batches are constructed through multiple time windows to provide a data basis for subsequent dependency analysis.
[0089] In step S320 , the standard deviation of each time series feature and the covariance between different time series features are calculated based on the time series feature group to analyze the mutual dependence between different time series features.
[0090] Here, the Pearson Correlation Coefficient (PCC) is used to measure the correlation between each feature to define the dependency between different features.
[0091]
[0092] Where Cov(f u ,f v ) is the time series feature f u With the time series characteristics f v The covariance between and Respectively represent f u The standard deviation and f v The standard deviation of uv represents f u With f v the degree of interdependence between them;
[0093] Here, standard deviation is used to measure the volatility of each feature in the sample set, reflecting whether the feature has significant changes. Features with large standard deviations indicate that their values change dramatically, may contain richer information, and should be given greater weight during the feature fusion process. Covariance is used to measure the common change trend of two features in the sample set. If two feature values increase or decrease at the same time, they have a positive correlation; if one increases and the other decreases, they have a negative correlation. Covariance can be used to identify features with strong correlations. The common changes of these features can be an important indicator of piracy behavior patterns. For example, if the update frequency and the frequency of functional module usage change in a coordinated manner, it may be a specific type of software piracy behavior.
[0094] Therefore, through the Pearson correlation coefficient, the potential relationship between temporal features and static features can be identified, so that those features with stronger correlation in piracy detection can be given higher weights.
[0095] In step S330 , a modified weight of each time series feature is calculated according to the correlation between different time series features.
[0096] Dynamically adjust the weight of each feature based on its dependency on other features. u The weight β u It can be determined by normalizing the weighted sum of its correlation with other features:
[0097]
[0098] Where, β u is the feature f u The weight of , S is the total number of features;
[0099] It should be noted that features in different time periods may have different influences. Dynamically adjusting weights allows the model to adapt to changes in temporal features, especially in pirated software detection scenarios with complex behavioral patterns. In this embodiment, unlike static weight assignment, a dynamic weighting mechanism is used by considering the interdependencies between features. This gives higher weights to features that are highly correlated with other features, thereby emphasizing their importance.
[0100] In step S340 , weighted correction is performed on the corresponding time series features based on the respective correction weights to obtain a time series feature matrix.
[0101] Finally, all features are fused by weighted summation to obtain the final fused time series feature matrix:
[0102]
[0103] Through the embodiments of the present application, by calculating the standard deviation and covariance, it is possible to identify specific behavior patterns within certain time periods. When multiple timing features (such as the frequency of use and update frequency of functional modules) change together in different time periods, it is possible to identify the patterns of interaction between these complex timing features, further improving the comprehensiveness of the detection. Thus, based on covariance and correlation analysis, the timing features are no longer simply weighted when fused, but by calculating their dependence on other features, the weight distribution in the fusion process is dynamically adjusted, ensuring the rationality of the timing features in the entire feature fusion and improving the responsiveness of the detection system to dynamic behaviors.
[0104] In some preferred implementations of the present application, the convolution kernel scale weights are adaptively determined using an entropy-based weighting method. Using entropy-based weighting, the model assigns weights based on the amount of information (i.e., entropy) in the feature matrix generated by each convolution kernel scale. This ensures that feature matrices with high information content receive higher weights, helping the model focus on more important features.
[0105] Figure 4 An operational flowchart of an example of adaptively determining convolution kernel scale weights based on information entropy weight calculation according to an embodiment of the present application is shown.
[0106] like Figure 4 As shown, in step S410, the eigenvalues in each scale feature matrix are calculated to construct a probability distribution.
[0107]
[0108] Where, Represents the feature matrix F i The eigenvalue of the jth element in , Represents eigenvalues The probability distribution of Represents eigenvalues Frequency of occurrence, is the feature matrix F i The total frequency of all eigenvalues in .
[0109] In step S420 , the information entropy of each scale feature matrix is calculated.
[0110]
[0111] Where, H(F i ) represents the feature matrix F i Information entropy.
[0112] In step S430, the convolution kernel scale weight corresponding to each convolution kernel is calculated:
[0113]
[0114] Where, α i Indicates the scale is k i ×k i The convolution kernel scale weights.
[0115] In this way, the weight of each convolution kernel scale is dynamically adjusted based on the information entropy of the feature matrix, ensuring that features with high information entropy account for a larger proportion of the final fused features. Feature matrices with high information entropy often contain more complex behavioral patterns or abnormal patterns. By assigning higher weights to these features, the model can more sensitively detect potential piracy. For example, during software usage, if a functional module is used at an unusually high frequency during certain time periods, the information entropy will capture the complexity of this pattern and assign a higher weight to the convolution kernel, increasing the model's sensitivity to this pattern. Furthermore, feature matrices with low information entropy often contain noise or irrelevant information. By automatically assigning lower weights, the model can effectively ignore irrelevant features and reduce the impact of noise on the model. For example, features that remain unchanged over a long period of time may be considered irrelevant behavior or noise, but these features will be assigned lower weights based on the information entropy calculation. Furthermore, the weight of each convolution kernel scale changes dynamically with the data, allowing the model to flexibly respond to feature changes in different scenarios and improving its ability to detect diverse behavioral patterns.
[0116] Furthermore, in some examples of the embodiments of the present application, the loss function of the piracy identification model adopts a hybrid loss function that includes classification loss and regularization loss based on information entropy. Specifically, the classification loss adopts the cross-entropy loss function, which is used to measure the difference between the predicted value and the true label in the classification task of the model, ensuring that the model can accurately classify pirated software and genuine software. In addition, by incorporating the information entropy regularization term, the model can dynamically adjust the weight of the convolution kernel scale according to the behavior of each sample during each training process to prevent over-reliance on certain high-weight feature dimensions. The regularization term ensures that the weight of the convolution kernel remains reasonably distributed, avoiding the excessive weight of some convolution kernels, which causes other feature dimensions to be ignored.
[0117] For example, the loss function of the piracy identification model is:
[0118]
[0119] Where, L total represents the loss function of the piracy identification model, represents the cross entropy loss of the mth sample, represents the information entropy regularization loss of the mth sample; is the actual category label of the mth sample, where the label value 0 represents a non-pirated software sample and the label value 1 represents a pirated software sample; is the model's predicted value of the piracy risk probability for the mth sample; C is the total number of classification categories, C = 2; M is the total number of data samples in the data sample set; is the mth sample in the convolution kernel scale k q ×k q The weight under , λ represents the loss balancing hyperparameter.
[0120] Here, λ can be a preset or learnable hyperparameter. By adjusting the hyperparameter λ, the model can flexibly balance between classification accuracy and feature weight distribution. For example, in scenarios where classification accuracy is more important, the impact of cross-entropy loss can be increased; in scenarios where feature weight distribution has a greater impact on performance, the role of the regularization term can be enhanced.
[0121] It should be noted that by introducing a regularization term based on information entropy, the model can dynamically adjust the weights of different convolution kernel scales according to the characteristics of each sample, so that the model can automatically identify which feature dimensions and time points are more important. This mechanism ensures that when the model faces complex behavior patterns (such as the frequency of software use, update response frequency, etc.), it can adaptively adjust the weights according to the performance of the features. For example, when detecting pirated software, specific behavior patterns (such as abnormal plug-in loading or abnormal frequency of use of functional modules) can be given higher weights, thereby improving recognition capabilities. At the same time, the weight adjustment process can dynamically reduce reliance on low-information or noise features, prevent the model from being misled by irrelevant features, and thus improve overall robustness and stability.
[0122] By using the loss function of the embodiment of the present application and cross entropy loss, the accuracy of the model in the classification task is ensured, that is, the model can accurately distinguish between pirated and genuine software. The information entropy regularization term is used to ensure the rationality of the weight distribution of the model in the multi-scale feature fusion process, avoiding the model's excessive dependence on certain features. As a result, the balance between classification performance and feature weight distribution can be flexibly adjusted. Through the balanced optimization of dual losses, each training iteration can simultaneously optimize the classification accuracy and feature weight distribution, find the optimal feature fusion solution more quickly, and enhance the adaptability of the model.
[0123] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of combined actions, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0124] Figure 5 A structural block diagram of an example of a genuine software detection and processing system based on big data analysis according to an embodiment of the present application is shown.
[0125] like Figure 5As shown, the genuine software detection and processing system 500 based on big data analysis includes an installation source acquisition unit 510, a piracy risk initial identification unit 520, a service log acquisition unit 530 and a piracy risk determination unit 540.
[0126] The installation source acquisition unit 510 is used to acquire the installation source information of each installed software in the client device to be detected.
[0127] The piracy risk initial identification unit 520 is used to verify the installation source information of each installation software based on the software installation authorization library to identify whether each installation software has potential piracy risks; the software authorization library contains multiple authorized genuine software names and corresponding software authorized installation channels.
[0128] The service log acquisition unit 530 is used to obtain the online service log of each risky installation software identified as having potential piracy risk, and extract multi-dimensional risk features from the online service log; the multi-dimensional risk features include the software online update response frequency, the software function module usage frequency, the software plug-in loading information and the software user behavior information.
[0129] The piracy risk determination unit 540 is used to input the multi-dimensional risk characteristics of each risky installed software into a piracy identification model to determine whether there is a piracy risk; the piracy identification model adopts a deep learning model.
[0130] In some embodiments, an embodiment of the present application provides a non-volatile computer-readable storage medium, which stores one or more programs including execution instructions, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute the steps of any of the above-mentioned methods for detecting and processing genuine software based on big data analysis in the present application.
[0131] In some embodiments, the embodiments of the present application also provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer performs any step of the above-mentioned genuine software detection and processing method based on big data analysis.
[0132] In some embodiments, an embodiment of the present application also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the genuine software detection and processing method based on big data analysis.
[0133] Figure 6 This is a hardware structure diagram of an electronic device that performs a genuine software detection and processing method based on big data analysis, provided in another embodiment of the present application. Figure 6 As shown, the device includes:
[0134] One or more processors 610 and memory 620, Figure 6 A processor 610 is taken as an example.
[0135] The device for executing the genuine software detection and processing method based on big data analysis may further include: an input device 630 and an output device 640 .
[0136] The processor 610, the memory 620, the input device 630 and the output device 640 may be connected via a bus or other means. Figure 6 The bus connection is taken as an example.
[0137] Memory 620, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the legitimate software detection and processing method based on big data analysis in the embodiments of this application. Processor 610 executes the non-volatile software programs, instructions, and modules stored in memory 620 to execute various server functions and data processing, thereby implementing the legitimate software detection and processing method based on big data analysis in the aforementioned method embodiment.
[0138] The memory 620 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 620 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 620 may optionally include a memory remotely located relative to the processor 610, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0139] The input device 630 may receive input digital or character information and generate signals related to user settings and function control of the electronic device. The output device 640 may include a display device such as a display screen.
[0140] The one or more modules are stored in the memory 620 and, when executed by the one or more processors 610, perform the genuine software detection and processing method based on big data analysis in any of the above method embodiments.
[0141] The above-mentioned product can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of this application.
[0142] The electronic devices of the embodiments of the present application exist in various forms, including but not limited to:
[0143] (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and their primary purpose is to provide voice and data communications. These terminals include smartphones, multimedia phones, feature phones, and low-end phones.
[0144] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers and have computing and processing capabilities, and generally also have mobile Internet access. These terminals include PDAs, MIDs, and UMPCs.
[0145] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players, handheld game consoles, e-books, smart toys, and portable car navigation devices.
[0146] (4) Other onboard electronic devices with data interaction functions, such as onboard computer devices installed in vehicles.
[0147] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0148] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiment.
[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for detecting and processing genuine software based on big data analysis, comprising: Obtaining the installation source information of each installed software in the client device to be detected; Verifying the installation source information of each installation software based on the software installation authorization library to identify whether each installation software has a potential piracy risk; For each risky installation software identified as having a potential piracy risk, obtaining an online service log of the risky installation software, and extracting a multi-dimensional risk feature from the online service log; The multi-dimensional risk characteristics include software online update response frequency, software function module usage frequency, software plug-in loading information, and software user behavior information; the online update response frequency is used to determine whether the software is frequently updated and whether it is consistent with the update cycle of genuine software; The functional module usage frequency is used to analyze the usage frequency of the software's core functions to identify whether it matches the normal usage behavior of genuine software; the plug-in loading information is used to analyze whether the software has loaded unauthorized plug-ins or extension modules; the software user behavior information is used to record user operation behavior to evaluate whether there are abnormal operations or abnormal usage patterns; Inputting the multi-dimensional risk characteristics of each of the risky installed software into a piracy identification model to determine whether there is a piracy risk; The piracy identification model includes: The CNN module is used to extract the local dependency between the software online update response frequency and the software functional module usage frequency in the multi-dimensional risk feature to generate corresponding temporal pattern features; The MLP module is used to extract the nonlinear relationship between the software plug-in loading information and the software user behavior information in the multidimensional risk feature to generate corresponding static pattern features; The feature fusion module is used to fuse the temporal pattern features and the static pattern features to generate corresponding fusion features; The classification module is used to classify the fused features through a fully connected layer to output a risk probability value that the risky installed software is pirated software; The extraction of the time series feature matrix includes: Performing time series analysis on the online service log according to a preset time window size to extract a time series feature group, wherein the time series feature group includes time series features corresponding to a preset number of time windows, and the time series features include software online update response frequency and software function module usage frequency; Based on the time series feature group, the standard deviation of each time series feature and the covariance between different time series features are calculated to analyze the mutual dependence between different time series features: Where Cov(f u ,f v ) is the time series feature f u With the time series characteristics f v The covariance between and Respectively represent f u The standard deviation and f v The standard deviation of uv represents f u With f v the degree of interdependence between them; Calculate the corrected weight of each time series feature based on the correlation between different time series features: Where, β u is the feature f u The weight of , S is the total number of features; Based on each of the correction weights, the corresponding time series features are weighted and corrected to obtain the time series feature matrix X t :
2. The method according to claim 1, wherein The CNN module uses a multi-scale convolutional network module; the multi-scale convolutional network module introduces multiple convolution kernels of different scales to simultaneously capture short-term fine-grained and long-term coarse-grained feature behavior patterns: Where, F i is through a size of k i ×k i The feature matrix generated by the convolution kernel is Indicates the scale is k i ×k i The convolution kernel weight matrix, X t Represents the input time series feature matrix, Conv(·) represents the convolution process; N is the total number of convolution kernels, α i is of scale k i ×k i The convolution kernel scale weight corresponding to the convolution kernel.
3. The method according to claim 1, wherein The feature fusion module adopts a feature fusion module based on a self-attention mechanism and is used to generate fused features by performing the following operations: Receive the temporal pattern features and static pattern features corresponding to the first risky installed software; the feature matrix of the temporal pattern features is T is the total number of time steps, F1 is the dimension of the time series feature; the feature matrix of the static pattern feature is F2 is the dimension of static features; The temporal pattern features and static pattern features are projected separately through the linear transformation matrix to generate the corresponding query vector, key vector and value vector: Where, is the linear transformation weight matrix of query, key, and value of time series features; is the linear transformation weight matrix of the query, key, and value of the static feature, d is the dimension of the latent feature, which is used to unify the representation of all input features; Q s ,K s ,V s Represent the query vector, key vector and value vector corresponding to the time series features, Q p ,K p ,V p Represent the query vector, key vector and value vector corresponding to the static features respectively; For the interaction between temporal features and static features, attention calculations are performed separately: Z s =Attention(Q s ,K s ,V s ), Z p =Attention(Q p ,K p ,V p ), Where Attention(·) represents the dot product attention calculation mechanism; is a scaling factor to prevent the dot product value from being too large and affecting the gradient calculation; softmax represents the softmax activation function, which converts the correlation value into a weight coefficient through the Softmax activation operation. The weight coefficient is used to perform a weighted summation on V to generate the fused feature representation; K T represents the transpose of the key vector K; Z s is the temporal feature self-attention representation; Z p It is a static feature self-attention representation; By mixing the query vector and key vector of temporal features and static features, we can further capture the interaction between them: Z sp =Attention(Q s ,K p ,V p ), Z ps =Attention(Q p ,K s ,V s ), Where Z sp Represents the first interactive attention feature of temporal features to static features, Z ps The second interactive attention feature representing the static feature to the temporal feature; Fusing the temporal feature self-attention representation, the static feature self-attention representation, the first interactive attention feature, and the second interactive attention feature: WITH final =Z s +Z p +Z sp +Z ps , Where Z final Indicates the fusion feature corresponding to the first risky installed software.
4. The method according to claim 1, wherein The convolution kernel scale weight is adaptively determined by calculating the weight based on information entropy: Calculate the eigenvalues in each scale feature matrix to construct a probability distribution: Where, Represents the feature matrix F i The eigenvalue of the jth element in , Represents eigenvalues The probability distribution of Represents eigenvalues Frequency of occurrence, is the feature matrix F i The total frequency of all eigenvalues in ; Calculate the information entropy of each scale feature matrix: Where, H(F i ) represents the feature matrix F i Information entropy; Calculate the convolution kernel scale weight corresponding to each convolution kernel: Where, α i Indicates the scale is k i ×k i The convolution kernel scale weights.
5. The method according to claim 4, wherein The loss function of the piracy identification model is: Where, L total represents the loss function of the piracy identification model, represents the cross entropy loss of the mth sample, represents the information entropy regularization loss of the mth sample; is the actual category label of the mth sample, where the label value 0 represents a non-pirated software sample and the label value 1 represents a pirated software sample; is the model's predicted value of the piracy risk probability for the mth sample; C is the total number of classification categories, C = 2; M is the total number of data samples in the data sample set; is the mth sample in the convolution kernel scale k q ×k q The weight under , λ represents the loss balancing hyperparameter.
6. A system for detecting and processing genuine software based on big data analysis, for implementing the method according to any one of claims 1 to 5; the system comprising: An installation source acquisition unit, configured to acquire installation source information of each installed software in the client device to be detected; a piracy risk initial identification unit, configured to verify the installation source information of each installed software based on a software installation authorization library to identify whether each installed software has a potential piracy risk; a service log acquisition unit, configured to acquire, for each risky installation software identified as having a potential piracy risk, an online service log of the risky installation software, and extract multi-dimensional risk features from the online service log; The multi-dimensional risk characteristics include software online update response frequency, software function module usage frequency, software plug-in loading information and software user behavior information; a piracy risk determination unit, configured to input the multi-dimensional risk characteristics of each of the risky installed software into a piracy identification model to determine whether there is a piracy risk; The piracy identification model adopts a deep learning model.
Citation Information
Patent Citations
Risk terminal management method and device, storage medium and server
CN116186675A
Software pirate identification and tracking method and related equipment
CN118172072A