Risk Assessment Method for Data Dependence Relationships in Intelligent Mobile Application Development
The method constructs a data dependency network using Neo4j to assess risks in open-source model ecosystems, addressing the challenge of managing complex and dynamic data dependencies in smart mobile app development, enhancing ecosystem stability and reliability.
Patent Information
- Application Number
- CN202510517029.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing technology is unable to effectively manage and evaluate complex and dynamically changing data dependencies in smart mobile application development, making it difficult to identify and deal with data risks, affecting the credibility and stability of the open source model ecosystem.
Build a data dependency relationship network, use the Neo4j database for visualization and dynamic analysis, design trustworthiness risks and vulnerability quantitative indicators, and comprehensively evaluate the risks of data and models.
It improves the understanding and monitoring capabilities of dependent network structures, timely discovers potential vulnerable nodes, and improves the sustainability and stability of the open source model ecosystem.
Smart Images

Figure CN120068091B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a method for risk assessment of data dependency relationships in the development of intelligent mobile applications. Background Art
[0002] At present, with the rapid development of mobile Internet technology and the booming intelligent mobile application market, open-source models, with their characteristics such as high efficiency and flexibility, are deeply embedded in all aspects of the development of intelligent mobile applications, greatly promoting the enrichment of functions and the improvement of performance of intelligent mobile applications. From the personalized recommendation of social intelligent mobile applications to the risk assessment of financial intelligent mobile applications, open-source models play an irreplaceable and crucial role. However, with the wide popularization of open-source models in the field of intelligent mobile application development, the data dependency problems in their ecosystem have gradually emerged, becoming a thorny problem hindering the further development of this field and attracting high attention from intelligent mobile application developers and related enterprises.
[0003] The open-source model ecosystem in the field of intelligent mobile application development shows remarkable complexity. In terms of model types, they are rich and diverse and have their own focuses. Taking the application of computer vision technology in intelligent mobile applications as an example, in the development of image recognition intelligent mobile applications (such as shopping applications based on picture search for goods, intelligent photo album classification applications), open-source versions of the convolutional neural network (CNN) model architecture are often used. Among them, the classic AlexNet model can quickly achieve basic image classification functions and has been widely used in some early simple image recognition intelligent mobile applications; while the VGG series models, with their deeper network structures, perform excellently in the accuracy of image feature extraction and are suitable for intelligent mobile application scenarios with higher requirements for image recognition accuracy. In the scenario of natural language processing for intelligent mobile applications, such as intelligent voice assistant intelligent mobile applications, open-source implementations of the Transformer model architecture, such as GPT - Neo, etc., provide strong support for achieving smooth human-computer conversations and accurate semantic understanding. These different open-source models have significant differences in algorithm principles, network structures, and parameter settings, and each adapts to diverse functional requirements in intelligent mobile applications.
[0004] In terms of data dependencies, the development of intelligent mobile applications is highly dependent on external data sets. Take the development of a location-based mobile social application as an example. To implement functions such as accurate user interest recommendations and nearby friend matching, a large amount of user behavior data, location data, and social relationship data are required. These data come from a wide range of sources, which may include user profile data provided by third-party data suppliers, real-time location data obtained through the built-in location service interface of intelligent mobile applications, and social interaction data generated and uploaded by users during use. However, numerous problems follow. First, it is difficult to ensure the clarity of data sources. The data collection channels of some third-party data suppliers may have gray areas. For example, obtaining data beyond a reasonable range by inducing users to authorize. In the event of a data leak or infringement dispute, intelligent mobile application developers will face huge risks. Second, the compliance of data authorization is ambiguous. The scope of data usage rights and authorization periods for different data sources lack clear definitions. For example, the use of certain location data may only be authorized for specific functional modules. In subsequent function expansions of the application, if the usage scope is expanded without re-authorization, it will violate relevant laws and regulations. Third, the quality of data is uneven. In user behavior data, due to the diversity of user operation habits and differences in data collection devices, it is difficult to guarantee the accuracy and integrity of the data. For example, some users may fill in personal interest tags randomly, resulting in a large amount of data noise and affecting the accuracy of the recommendation model trained based on this. These data problems directly affect the performance and stability of the open-source model in the actual operation of intelligent mobile applications, and further reduce the credibility of the entire open-source model ecosystem in the field of intelligent mobile application development.
[0005] Open-source model platforms play a central hub role in the field of intelligent mobile application development. Taking Hugging Face as an example, it provides comprehensive technical support, rich model resources, and efficient model hosting services for the open-source community of intelligent mobile applications. As more and more open-source models for intelligent mobile application development are released and widely used on this platform, Hugging Face has become a key node for model sharing and storage in the field of intelligent mobile applications. With the explosion in the number of intelligent mobile applications and the increasing demand for personalized features from users, activities such as model reuse, adaptation, and citation are becoming increasingly frequent. In the development of mobile e-commerce applications, developers often conduct secondary development based on existing open-source recommendation models, such as open-source implementations of collaborative filtering algorithms, to adapt to the unique user behavior data and product attribute data of the mobile e-commerce platform, achieve accurate product recommendations, and increase the user purchase conversion rate. In mobile game development, enterprises will cite open-source artificial intelligence models, such as open-source versions of reinforcement learning models for game character behavior decision-making, and optimize them in combination with in-game player behavior data and game scenario data to enhance the fun and challenge of the game. Against this backdrop, how to efficiently track and manage data dependencies in the unique processes of intelligent mobile application development, such as rapid iterative development, multi-platform adaptation, and complex project management environments, has become an urgent challenge for the open-source model ecosystem in the field of intelligent mobile application development.
[0006] The model building process in intelligent mobile application development encompasses multiple key stages, including data collection, cleaning, annotation, model pre-training, fine-tuning, evaluation, and integration into intelligent mobile application projects. In these stages, data, models, and related APIs are closely associated and interdependent. Taking the application of machine learning models in the development of mobile image editing applications as an example, in the data collection stage, a large amount of image data with different styles and resolutions needs to be collected, and the quality of this data directly affects the effect of subsequent model training. In the data cleaning stage, blurred and damaged image data needs to be removed to ensure the usability of the data. In the annotation stage, professionals are required to accurately annotate the images, such as marking the areas that need to be processed with special effects. In the model pre-training stage, a large-scale general image dataset is used for preliminary training on a selected open-source model architecture. In the fine-tuning stage, according to the specific requirements of the mobile image editing application, such as optimizing the image special effect for the mobile phone screen resolution, the pre-trained model is adjusted using the in-app user-generated image data. In the evaluation stage, a series of metrics, such as the accuracy and processing speed of image special effect processing, are used to judge the model performance. Finally, in the integration stage, the model needs to interact with other modules of the intelligent mobile application, such as the user interface interaction module, the image storage module, and the underlying mobile operating system, through specific APIs to achieve seamless docking of functions. This complex dependency greatly increases the complexity of the open-source model ecosystem in the field of intelligent mobile application development. For researchers in the field of intelligent mobile application development, a deep understanding of these complex data flows and dependencies helps to optimize the operating efficiency of the model in the limited resource environment of mobile devices and improve the adaptability of the model to the specific functional requirements of intelligent mobile applications; for intelligent mobile application developers, a transparent and traceable model building process can enhance the credibility and usability of the model in intelligent mobile application projects, making them more confident in applying the model to core function development, such as using the model for security risk assessment in mobile payment applications; for data owners, ensuring the compliance of data in the intelligent mobile application development environment, such as strictly following the security specifications for mobile device data storage and transmission for license management to prevent the theft or abuse of user data during mobile network transmission, is an important prerequisite for responsible data use. Therefore, formulating effective strategies to manage and track data dependencies is of great significance for ensuring the sustainable development and reliability of the open-source model ecosystem in the field of intelligent mobile application development.
[0007] Currently, there are obvious deficiencies in the research on data dependency relationships in the open-source model ecosystem in the field of intelligent mobile application development. There is a lack of both a systematic and comprehensive macro perspective and practical and effective measurement tools. When dealing with data dependency problems in intelligent mobile application development, traditional software package dependency analysis methods are unable to cope with complex data dependency structures and ever-changing characteristics. Traditional methods are mainly applicable to the tracking of static dependency relationships and usually can only handle limited types of dependencies, making it difficult to deal with complex scenarios involving multiple intelligent mobile application models and datasets with continuously dynamic dependency relationships. For example, during the frequent update and iteration of intelligent mobile applications, user behavior data patterns change continuously, and new functional requirements lead to changes in dependency relationships with different types of datasets. Traditional methods cannot respond promptly and effectively. Traditional software package dependency analysis methods mainly target dependencies between program codes or software packages and cannot adapt to the data-driven model development environment in the field of intelligent mobile application development. The open-source model ecosystem in the field of intelligent mobile application development is far from simple software package interdependencies and also involves multiple complex aspects such as the management of data storage in mobile devices, transmission through mobile networks, and the interaction between model training and the rapid development cycle of intelligent mobile applications. This complexity requires dependency analysis methods to have richer levels and dimensions, while existing methods are difficult to meet such requirements. For example, during the fine-tuning stage of intelligent mobile application models, there are essential differences in logic between data dependencies and version dependencies of code libraries, and existing methods cannot effectively distinguish and handle them properly. Currently, most intelligent mobile application model ecological risk assessment methods focus on qualitative analysis and lack quantitative criteria and operable indicators. Traditional compliance analysis methods mainly focus on the source and authorization of data but fail to extract quantifiable risk points from the full-process dependency relationships in the construction of intelligent mobile application models. For example, they do not fully consider the compliance differences in data usage of intelligent mobile applications in different regions and different network environments, as well as the potential risks during data caching and local storage in mobile devices. Therefore, existing technologies cannot provide a clear and operable risk measurement system, making it difficult to assist users in the field of intelligent mobile application development to effectively identify and respond to potential data risks. Summary of the Invention
[0008] Based on this, in view of the above technical problems, it is necessary to provide a risk assessment method for data dependency relationships in intelligent mobile application development that can assist users in the field of intelligent mobile application development to effectively identify and respond to potential data risks.
[0009] A risk assessment method for data dependency relationships in intelligent mobile application development, the method comprising:
[0010] Obtain data from the HF platform to construct a dataset; preprocess the dataset to obtain a preprocessed dataset; define six entities according to the preprocessed dataset, including models, datasets, developers, annotators, tasks, and papers; construct a data dependency network based on the six entities and the key attributes between the entities;
[0011] Build a visualization tool based on the Neo4j database to identify the structure and evolution information in the data dependency network through relationship queries, developer queries, and node path queries; the structure and evolution information include vulnerable nodes and risk paths;
[0012] Design quantitative metrics using the structure and evolution information to obtain credibility risk metrics and vulnerability metrics; conduct a risk assessment of the data dependency relationship based on the credibility risk metrics and vulnerability metrics.
[0013] In the above data dependency relationship risk assessment method for intelligent mobile application development, this application constructs a comprehensive data dependency network, which can comprehensively analyze the dependency relationships between various components in the open-source model ecosystem. By using the Neo4j database to manage complex dependencies and combining dynamic evolution analysis, it can accurately capture the changing trends of data and model dependencies. This method solves the problem that traditional methods cannot effectively handle large-scale and dynamically changing dependency relationships, and greatly improves the ability to understand and monitor the structure of the dependency network. At the same time, "credibility risk" and "vulnerability" quantitative metrics are designed to make the risk assessment more accurate and intuitive. These two new metrics quantitatively evaluate the risks of data and models from multiple dimensions such as attribute integrity, timeliness, community feedback, and sensitivity to external dependencies, helping managers promptly discover potential vulnerable nodes or untrusted dependencies, avoiding the defect that traditional methods cannot comprehensively evaluate risks, and thus enhancing the sustainability and stability of the open-source model ecosystem. Through the innovative combination of quantitative evaluation and dynamic analysis, the deficiencies in data dependency relationship management and risk assessment in the existing technology are solved. It not only improves the stability and reliability of the system, but also provides an effective technical means for the healthy management of the open-source model ecosystem for intelligent mobile application development, and can help users in the field of intelligent mobile application development effectively identify and respond to potential data risks. Brief Description of the Drawings
[0014] Figure 1 It is a schematic flowchart of a data dependency relationship risk assessment method for intelligent mobile application development in an embodiment;
[0015] Figure 2 It is a schematic diagram of a data dependency network in an embodiment;
[0016] Figure 3 It is an internal structure diagram of a computer device in an embodiment. Detailed Implementation Manner
[0017] In order to make the objectives, technical solutions, and advantages of this application clearer and more understandable, the following further details this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not used to limit this application.
[0018] In one embodiment, as Figure 1 shown, a method for risk assessment of data dependency relationships in intelligent mobile application development is provided, including the following steps:
[0019] Step 102, obtain the data construction dataset of the HF platform; preprocess the dataset to obtain the preprocessed dataset; define six types of entities according to the preprocessed dataset, including models, datasets, developers, annotators, tasks, and papers; construct a data dependency network based on the six types of entities and the key attributes between the entities.
[0020] By obtaining the data construction dataset of the HF (Hugging Face) platform, the data from the HF (Hugging Face) platform, which hosts more than 400,000 AI models and 100,000 datasets, has frequent data updates and rich resources. Data from the HF community is selected, and a data dependency network is constructed to comprehensively capture the dependency relationships and their evolution characteristics among various components in the open-source model ecosystem.
[0021] After preprocessing the dataset, six types of entities are defined, and a data dependency network is constructed based on these entities and their key attributes. Data dependencies usually refer to models or datasets that declare some data information (such as the author of the model, the license of the model, etc., the author of the dataset, the license of the dataset, the dataset of the model, the base model of the model), etc. On the one hand, due to the dependency relationships, and on the other hand, due to the data, complex data dependency relationships are formed. This enables the various components and their dependency relationships in the originally complex and disorderly open-source model ecosystem to be presented in a structured manner. Compared with traditional methods, it more comprehensively covers the associations between data, models, etc. in intelligent mobile application development, thereby effectively managing complex data dependency relationships. In the development of mobile image editing applications, the relationships between the model and the image dataset used for training, the improvements made by developers to the model, the annotations of data by annotators, etc. can all be clearly reflected in this data dependency network, facilitating developers to understand the dependency context throughout the development process.
[0022] Step 104, construct a visualization tool based on the Neo4j database to identify the structure and evolution information in the data dependency network through relationship queries, developer queries, and node path queries; the structure and evolution information include vulnerable nodes and risk paths.
[0023] Neo4j is selected as the underlying database management platform. Neo4j can effectively manage complex data relationships, which makes it very suitable for describing the dependency structure in the model ecosystem. The overall architecture of the data dependency network is as follows Figure 2 shown. When constructing the data dependency network, this application defines six main entities: "Model", "Dataset", "Author", i.e., the developer, "Annotator", "Task", and "Paper". Each entity has multiple key attributes. For example, the author, creation time, base model, and associated datasets of the model; the dataset contains information such as creation time, download times, annotator, and license. These attributes come from the tags on the HF platform and are mainly extracted from the model card and dataset card, forming the basis of the data dependency network. To deeply analyze the evolution of the data dependency network, the network is divided into multiple time windows, and based on the creation timestamps of the entities, it is observed how the dependencies change over time. This process not only reveals the deepening dependency relationship between models and datasets but also reflects the continuous reuse and interaction of datasets and models. By analyzing the network structure within different time windows, the dynamic evolution characteristics of the dependency relationship can be deeply understood.
[0024] According to the label description, the relationships between datasets are divided into two types: "SOURCE DATAFROM" and "EXTEND DATA FROM". The former points to the original dataset, while the latter indicates that the dataset is an extension or derivative based on other datasets. Through this label information, this application can capture the full picture of data dependencies, thus providing support for subsequent risk assessment.
[0025] Based on the Neo4j database, a visualization tool is built. Through methods such as relationship query, developer query, and node path query, it can identify the structure and evolution information in the data dependency network, including vulnerable nodes and risk paths. The Neo4j database is good at processing complex graph-structured data and can efficiently manage the many-to-many and dynamically changing dependency relationships that exist in large numbers in the open-source model ecosystem. The visualization tool allows developers to more intuitively see this information and timely discover potential problems in the dependency relationship.
[0026] During the frequent update and iteration of intelligent mobile applications, when the change in the user behavior data pattern leads to a change in the dependency relationship on the dataset, the visualization tool can quickly locate the affected nodes and paths, so as to timely adjust the development strategy, solving the problem that traditional methods cannot respond to dynamic changes in a timely manner.
[0027] Step 106: Design quantization metrics using structural and evolutionary information to obtain credibility risk metrics and vulnerability metrics; conduct data dependency relationship risk assessment based on the credibility risk metrics and vulnerability metrics.
[0028] The "credibility risk" and "vulnerability" quantization metrics designed in this application. These two metrics quantify and evaluate the risks of data and models from multiple dimensions such as attribute integrity (such as model author information, license integrity), timeliness (the impact of the time after the model or dataset is released on its credibility and vulnerability), community feedback (the number of "likes" per day on average reflects community recognition), and sensitivity of external dependencies (the impact of the node's dependency relationship on upstream nodes on vulnerability). It focuses on evaluating the subjective risks of the system (i.e., the autonomy and independence of the system when facing the external environment) and objective risks (i.e., the vulnerability of the system when affected by external changes).
[0029] Different from traditional methods that focus on qualitative analysis and only pay attention to the compliance analysis of data sources and authorizations, these quantization metrics can extract quantifiable risk points from the full-process dependency relationship of the intelligent mobile application model construction, making the risk assessment more accurate and intuitive. When using the model for security risk assessment in mobile payment applications, the "credibility risk" metric can evaluate the credibility of the model itself, and the "vulnerability" metric can evaluate the stability of the model when external dependencies change, helping managers timely discover potential vulnerable nodes or untrustworthy dependencies, and effectively identify and respond to potential data risks.
[0030] In the above data dependency risk assessment method for intelligent mobile application development, this application can comprehensively analyze the dependency relationships between various components in the open-source model ecosystem by constructing a comprehensive data dependency network. By using the Neo4j database to manage complex dependency relationships and combining dynamic evolution analysis, it can accurately capture the changing trends of data and model dependency relationships. This approach solves the problem in traditional methods of being unable to effectively handle large-scale and dynamically changing dependency relationships, greatly improving the understanding and monitoring capabilities of the dependency network structure. At the same time, "credibility risk" and "vulnerability" quantification indicators are designed to make risk assessment more accurate and intuitive. These two new indicators quantitatively evaluate the risks of data and models from multiple dimensions such as attribute integrity, timeliness, community feedback, and sensitivity of external dependencies, helping managers promptly discover potential vulnerable nodes or untrusted dependencies, avoiding the defect of traditional methods being unable to comprehensively assess risks, and thus enhancing the sustainability and stability of the open-source model ecosystem. Through the innovative combination of quantitative evaluation and dynamic analysis, the deficiencies in data dependency relationship management and risk assessment in the existing technology are solved. It not only improves the stability and reliability of the system but also provides an effective technical means for the healthy management of the open-source model ecosystem for intelligent mobile application development, enabling users in the field of intelligent mobile application development to effectively identify and respond to potential data risks.
[0031] In one embodiment, the data set is preprocessed to obtain a preprocessed data set, including:
[0032] After converting the CreatedAt field in the data set to the date and time format and extracting the developer information, Space type data and sparse or irrelevant fields, including SHA hash values and redundant tags, are removed to obtain the preprocessed data set.
[0033] In a specific embodiment, during the data collection process, the HFCommunity data set is used, which covers three types of repositories on the HF platform: models, data sets, and spaces. To improve the accuracy of data analysis, the following steps are taken when processing the data:
[0034] Convert the "CreatedAt" field to the date and time format; extract the developer information; and remove sparse or irrelevant fields such as SHA hash values and redundant tags (such as "tag-TikTok" and "tag-Twitter").
[0035] Since repositories of the "Space" type mainly serve as examples of user interfaces, application displays, and model interactions and do not make a significant contribution to the analysis of data dependency relationships in the model ecosystem, these data are excluded.
[0036] In one embodiment, quantization metrics are designed using structural and evolutionary information to obtain credibility risk metrics and vulnerability metrics, including:
[0037] Based on the structural and evolutionary information, the credibility risk of data nodes is evaluated from three aspects: attribute integrity, timeliness, and the user perspective, and the credibility risk metric is constructed by combining the exponential function and the logarithmic function.
[0038] In a specific embodiment, for a given data node, its "credibility risk" is mainly evaluated from three aspects:
[0039] Attribute integrity. Node attributes (such as author, license) and relationship attributes (such as source dataset or model) are crucial for ensuring legality and transparency. The absence of attributes such as license information may affect compliance. For almost isolated nodes (with only one inbound or outbound edge), the intrinsic attributes are the main concern. For nodes in a dependency chain, the risks associated with incomplete attributes may propagate through the entire chain, affecting the overall reliability.
[0040] Timeliness. After a model or dataset is released, the corresponding real-world scenarios are often dynamic, which may lead to concept drift or data bias. Therefore, timeliness is crucial for the "credible risk" of models and datasets.
[0041] From the user's perspective, nodes with higher content and more "likes" are usually subject to more frequent reviews and verifications, especially within high-frequency communities. User feedback can improve the file quality and increase opportunities for improvement.
[0042] When constructing the "credibility risk" metric, the exponential function and the logarithmic function are combined to effectively capture key risk factors and their interactions. The exponential function amplifies the impact of missing attributes, such as the absence of author information or license, because these absences seriously undermine the legality and credibility of the node. Therefore, a non-linear method is needed to accurately represent this potential threat. The logarithmic function is used to describe the gradually accumulating risks, such as time-related effects, so as to capture the gradual changes within the system and avoid a sudden drop in credibility due to information obsolescence. The average number of "likes" per day is also incorporated into the risk metric to comprehensively reflect the community's recognition. Relying solely on the number of "likes" may lead to biased evaluations, especially being disadvantageous to newer repositories due to time accumulation. By combining "likes" with "time", it is possible to better balance popularity and timeliness and reduce biases related to time differences. The integration of factors includes multiplication and division: multiplication captures the cumulative effects of different risk factors, while division helps control the overall risk level and prevent linear or excessive growth of risks. In addition, the number of subordinate nodes is added to the denominator to reflect the supervisory role of downstream nodes.
[0043] In one embodiment, the credibility risk indicator is:
[0044] ;
[0045] Wherein, represents the developer of the data node, and the developer is part of the node attributes, represents the license, and the license information is also one of the node attributes, represents the time interval, reflecting the time span from the release of the model or dataset to the current time, represents the average number of "likes" per day, reflecting the community's recognition, represents the weight coefficient related to the node attributes, represents the time base value.
[0046] In one embodiment, according to the outbound relationship of the node in the structure and evolution information, that is, the dependency relationship of the node on the upstream node, and the credibility risk of the node itself, the vulnerability indicator is obtained through relevant function mapping.
[0047] In a specific embodiment, the "vulnerability" indicator is different from traditional software. The situation in the model field is different: on the one hand, multiple dependency relationships can complement and enhance each other. For example, there may be many data sources with uneven quality, but this diversity helps to reduce the risk of bias caused by incomplete datasets. On the other hand, some data dependencies are one-time, such as datasets only used in the training phase. This means that "vulnerability" is not only the cumulative effect of upstream risks, but also closely related to the intrinsic quality of the node itself.
[0048] To effectively address these unique challenges of artificial intelligence models, this application considers two key aspects.
[0049] (i) Considered the outbound relationship of the node, which represents the dependency relationship of the node on the upstream node. The more outbound links a node has, the more dependency relationships it maintains, which can increase or decrease its "vulnerability", depending on the complementarity of these dependency relationships. The importance of each relationship is further distinguished using the weighting factor W.
[0050] (ii) Different from traditional software measurement methods, this application emphasizes the credibility risk of the node itself. A node with a lower credibility risk indicates better data integrity, timeliness, and higher community recognition, so the "vulnerability" is lower. In addition, the "credibility risk" indicator is also used to balance the impact of upstream nodes on "vulnerability", and the internal quality of the node can offset part of the negative impact of external dependencies.
[0051] In one embodiment, the vulnerability indicator is:
[0052] ;
[0053] wherein, represents the number of upstream nodes pointed to by the node, and the upstream nodes represent the relevant nodes in the outbound relationship, represents the th upstream pointed to by the node, i and is the credibility risk index value of the upstream node, represents the credibility risk index value of the current node , represents the belief mapping function that affects the vulnerability of the credibility of the current node, represents the weight of the relationship connecting the upstream node and the current node , represents the upstream node pointed to by the representative node , that is, the node with an outbound relationship, represents the node for which the vulnerability index needs to be calculated currently.
[0054] In one embodiment, relationship queries are used to explore the graph structure of a specific type of relationship; developer queries are used to visualize a specific developer node and its relationships; node path queries are used to show the shortest path between nodes.
[0055] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0056] at least a part of the steps in Figure 3As shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements a method for risk assessment of data dependency relationships in intelligent mobile application development. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0057] Those skilled in the art can understand that Figure 3 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0058] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0059] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0060] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for risk assessment of data dependency relationships in intelligent mobile application development, characterized in that The method includes: Obtaining data from the Hugging Face platform to construct a dataset; preprocessing the dataset to obtain a preprocessed dataset; defining six entities according to the preprocessed dataset, including models, datasets, developers, annotators, tasks, and papers; constructing a data dependency network based on the six entities and the key attributes between the entities. Constructing a visualization tool based on the Neo4j database to identify the structure and evolution information in the data dependency network through relationship queries, developer queries, and node path queries; the structure and evolution information including vulnerable nodes and risk paths. Designing quantitative metrics using the structure and evolution information to obtain credibility risk metrics and vulnerability metrics; performing a risk assessment of data dependency relationships based on the credibility risk metrics and vulnerability metrics.
2. The method according to claim 1, characterized in that, Preprocessing the dataset to obtain a preprocessed dataset, including: Converting the CreatedAt field in the dataset to a datetime format, extracting developer information, and then removing Space type data and sparse or irrelevant fields, including SHA hash values and redundant tags, to obtain a preprocessed dataset.
3. The method according to claim 1, wherein Designing quantitative metrics using the structure and evolution information to obtain credibility risk metrics and vulnerability metrics, including: Evaluating the credibility risk of data nodes from three aspects: attribute integrity, timeliness, and user perspective according to the structure and evolution information, and constructing credibility risk metrics by combining exponential functions and logarithmic functions.
4. The method according to claim 3, wherein The credibility risk metric is: ; Among them, represents the developer of the data node, and the developer is part of the node attributes, represents the license, and the license information is also one of the node attributes, represents the time interval, reflecting the time span from the release of the model or dataset to the present, represents the average number of "likes" per day, reflecting the recognition of the community, represents the weight coefficient related to the node attributes, represents the time reference value.
5. The method according to claim 3, characterized in that, The method further includes: Obtaining the vulnerability metric by mapping through relevant functions according to the outbound relationships of nodes in the structure and evolution information, that is, the dependency relationships of nodes on upstream nodes, and the credibility risk of the nodes themselves.
6. The method according to claim 5, characterized in that, The vulnerability metric is: ; Among them, represents the number of upstream nodes pointed to by the node and represents the th i upstream credibility risk index value represents the credibility risk index value of the current node ; represents the belief mapping function that affects the vulnerability of the credibility of the current node, represents the weight of the relationship between the upstream node and the current node ; represents the upstream node pointed to by the representative node , that is, the node with an outbound relationship represents the node for which the vulnerability index needs to be calculated currently.
7. The method according to claim 1, characterized in that, The relationship query is used to explore the graph structure of specific types of relationships; the developer query is used to visualize a specific developer node and its relationships; the node path query is used to display the shortest path between nodes.
Citation Information
Patent Citations
Container cluster risk analysis and vulnerability assessment method and device based on environmental perception
CN117874768A
Multi-source software supply chain intelligent analysis method and system
CN119720225A