Data dependency risk assessment method in intelligent mobile application development

By building data dependency networks and designing quantitative indicators, the problem of complex and dynamic changes in data dependencies in the field of intelligent mobile application development is solved, efficient risk assessment and management is achieved, and the sustainability and stability of the system are improved.

CN120068091AActive Publication Date: 2025-05-30NAT UNIV OF DEFENSE TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510517029.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-30
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

In the open source model ecosystem in the field of intelligent mobile application development, data dependencies are complex and dynamically changing, and it is difficult for existing technologies to effectively manage and evaluate data risks, resulting in the impact of the sustainability and stability of the system.

Method used

By building a data dependency network, using Neo4j database to manage complex dependencies, and combining dynamic evolution analysis, "credibility risk" and "frailty" quantitative indicators were designed to conduct data dependency risk assessment.

Benefits of technology

It realizes accurate capture of the change trends of data and model dependency relationships, improves the understanding and monitoring of the dependency network structure, enhances the accuracy and intuitiveness of risk assessment, and improves the sustainability and stability of the open source model ecosystem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068091A_ABST
    Figure CN120068091A_ABST
Patent Text Reader

Abstract

The invention relates to a data dependency risk assessment method in intelligent mobile application development. The method comprises the following steps: acquiring data of an HF platform to construct a data set; preprocessing the data set to obtain a preprocessed data set; defining six entities according to the preprocessed data set, wherein the six entities comprise a model, a data set, a developer, an annotator, a task and a paper; constructing a data dependence network according to the six entities and the key attributes among the entities; constructing a visual tool based on a Neo4j database, and identifying structure and evolution information in the data dependence network through relation query, developer query and node path query; the structure and evolution information comprises fragile nodes and risk paths; and designing a quantitative index according to the structure and evolution information to obtain a credibility risk index and a vulnerability index. By adopting the method, users in the field of intelligent mobile application development can be assisted to effectively identify and cope with potential data risks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a method for risk assessment of data dependency relationships in the development of intelligent mobile applications. Background Art

[0002] At present, with the rapid development of mobile Internet technology and the booming intelligent mobile application market, open-source models, with their characteristics such as high efficiency and flexibility, are deeply embedded in all aspects of the development of intelligent mobile applications, greatly promoting the enrichment of functions and the improvement of performance of intelligent mobile applications. From the personalized recommendation of social intelligent mobile applications to the risk assessment of financial intelligent mobile applications, open-source models play an irreplaceable and crucial role. However, with the wide popularization of open-source models in the field of intelligent mobile application development, data dependency problems in its ecosystem have gradually emerged, becoming a thorny problem hindering the further development of this field and attracting high attention from intelligent mobile application developers and related enterprises. The open-source model ecosystem in the field of intelligent mobile application development presents remarkable complexity. In terms of model types, they are rich and diverse and have different focuses. Taking the application of computer vision technology in intelligent mobile applications as an example, in the development of image recognition intelligent mobile applications (such as shopping applications based on picture search for goods, intelligent photo album classification applications), open-source versions of the convolutional neural network (CNN) model architecture are often used. Among them, the classic AlexNet model can quickly implement basic image classification functions and has been widely used in some early simple image recognition intelligent mobile applications; while the VGG series models, with their deeper network structures, perform excellently in the accuracy of image feature extraction and are suitable for intelligent mobile application scenarios with higher requirements for image recognition accuracy. In the scenario of natural language processing for intelligent mobile applications, such as intelligent voice assistant intelligent mobile applications, open-source implementations of the Transformer model architecture, such as GPT - Neo, etc., provide strong support for achieving smooth human-machine dialogue and accurate semantic understanding. These different open-source models have significant differences in algorithm principles, network structures, and parameter settings, and are respectively adapted to diverse functional requirements in intelligent mobile applications. In terms of data dependencies, the development of intelligent mobile applications is highly dependent on external data sets. Taking the development of a location-based mobile social application as an example, to achieve functions such as accurate user interest recommendations and nearby friend matching, a large amount of user behavior data, location data, and social relationship data are required. These data come from a wide range of sources, which may include user portrait data provided by third-party data providers, real-time location data obtained through the built-in location service interface of intelligent mobile applications, and social interaction data generated and uploaded by users during use. However, numerous problems follow. First, it is difficult to ensure the clarity of data sources. The data collection channels of some third-party data providers may have gray areas. For example, obtaining data beyond a reasonable range by inducing users to grant authorization. Once a data leak or infringement dispute occurs, intelligent mobile application developers will face huge risks. Second, data authorization compliance is ambiguous. The scope of data usage rights and authorization periods for different data sources lack clear definitions. For example, the use of certain location data may only be authorized for specific functional modules. If the scope of use is expanded without re-authorization during subsequent function expansion of the application, it will violate relevant laws and regulations. Third, the quality of data is uneven. In user behavior data, due to the diversity of user operation habits and differences in data collection devices, it is difficult to guarantee the accuracy and integrity of the data. For example, some users may fill in personal interest tags randomly, resulting in a large amount of data noise and affecting the accuracy of the recommendation model trained based on this data. These data problems directly affect the performance and stability of the open-source model in the actual operation of intelligent mobile applications, thereby reducing the credibility of the entire open-source model ecosystem in the field of intelligent mobile application development. Open-source model platforms play a central hub role in the field of intelligent mobile application development. Taking Hugging Face as an example, it provides comprehensive technical support, rich model resources, and efficient model hosting services for the open-source community of intelligent mobile applications. As more and more open-source models for intelligent mobile application development are released and widely used on this platform, Hugging Face has become a key node for model sharing and storage in the field of intelligent mobile applications. With the rapid increase in the number of intelligent mobile applications and the growing demand for personalized features from users, activities such as model reuse, adaptation, and citation are becoming more frequent. In the development of mobile e-commerce applications, developers often conduct secondary development based on existing open-source recommendation models, such as open-source implementations of collaborative filtering algorithms, to adapt to the unique user behavior data and product attribute data of the mobile e-commerce platform, achieve accurate product recommendations, and improve the user purchase conversion rate. In mobile game development, enterprises will cite open-source artificial intelligence models, such as open-source versions of reinforcement learning models for game character behavior decision-making, and optimize them in combination with in-game player behavior data and game scenario data to enhance the fun and challenge of the game. In this context, how to efficiently track and manage data dependencies in the specific processes of intelligent mobile application development, such as rapid iterative development, multi-platform adaptation, and complex project management environments, has become an urgent challenge for the open-source model ecosystem in the field of intelligent mobile application development. The model building process in intelligent mobile application development encompasses multiple key stages, including data collection, cleaning, annotation, model pre-training, fine-tuning, evaluation, and integration into intelligent mobile application projects. In these stages, data, models, and related APIs are closely associated and interdependent. Taking the application of machine learning models in the development of mobile image editing applications as an example, in the data collection stage, a large amount of image data with different styles and resolutions needs to be collected, and the quality of this data directly affects the effect of subsequent model training. In the data cleaning stage, blurred and damaged image data needs to be removed to ensure the usability of the data. In the annotation stage, professionals are required to accurately annotate the images, such as marking the areas that need special effects processing. In the model pre-training stage, preliminary training is carried out on a selected open-source model architecture using a large-scale general image dataset. In the fine-tuning stage, according to the specific requirements of the mobile image editing application, such as optimizing the image special effects for the mobile phone screen resolution, the pre-trained model is adjusted using the in-app user-generated image data. In the evaluation stage, a series of metrics, such as the accuracy and processing speed of image special effects processing, are used to judge the model performance. Finally, in the integration stage, the model needs to interact with other modules of the intelligent mobile application, such as the user interface interaction module, the image storage module, and the underlying mobile operating system, through specific APIs to achieve seamless function docking. This complex dependency greatly increases the complexity of the open-source model ecosystem in the field of intelligent mobile application development. For researchers in the field of intelligent mobile application development, a deep understanding of these complex data flows and dependencies helps optimize the running efficiency of the model in the limited resource environment of mobile devices and improve the adaptability of the model to the specific functional requirements of intelligent mobile applications; for intelligent mobile application developers, a transparent and traceable model building process can enhance the credibility and usability of the model in intelligent mobile application projects, giving them more confidence to apply the model to core function development, such as using the model for security risk assessment in mobile payment applications; for data owners, ensuring the compliance of data in the intelligent mobile application development environment, such as strictly following the security specifications for mobile device data storage and transmission for license management to prevent user data from being stolen or misused during mobile network transmission, is an important prerequisite for responsible data use. Therefore, formulating effective strategies to manage and trace data dependencies is of great significance for ensuring the sustainable development and reliability of the open-source model ecosystem in the field of intelligent mobile application development. Currently, there are obvious deficiencies in the research on data dependency relationships in the open-source model ecosystem in the field of intelligent mobile application development. There is a lack of a systematic and comprehensive macro perspective, as well as practical and effective measurement tools. When dealing with data dependency problems in intelligent mobile application development, traditional software package dependency analysis methods are unable to cope with the complex data dependency structures and ever-changing characteristics. Traditional methods are mainly applicable to the tracking of static dependency relationships and usually can only handle limited types of dependencies, making it difficult to handle complex scenarios involving multiple intelligent mobile application models and data sets with continuously changing dependency relationships. For example, during the frequent update and iteration of intelligent mobile applications, the patterns of user behavior data change continuously, and new functional requirements lead to changes in dependency relationships with different types of data sets. Traditional methods cannot respond promptly and effectively. Traditional software package dependency analysis methods mainly target dependencies between program codes or software packages and cannot adapt to the data-driven model development environment in the field of intelligent mobile application development. The open-source model ecosystem in the field of intelligent mobile application development is far from simple software package interdependencies and also involves multiple complex aspects such as the management of data storage on mobile devices, transmission through mobile networks, and the interaction between model training and the rapid development cycle of intelligent mobile applications. This complexity requires dependency analysis methods to have richer levels and dimensions, while existing methods are difficult to meet such requirements. For example, during the fine-tuning stage of intelligent mobile application models, there are essential differences in logic between data dependencies and version dependencies of code libraries, and existing methods cannot effectively distinguish and handle them properly. Currently, most intelligent mobile application model ecological risk assessment methods focus on qualitative analysis and lack quantitative criteria and operable indicators. Traditional compliance analysis methods mainly focus on the source and authorization of data but fail to extract quantifiable risk points from the full-process dependency relationships in the construction of intelligent mobile application models. For example, they do not fully consider the compliance differences in data usage in different regions and different network environments for intelligent mobile applications, as well as the potential risks during data caching and local storage on mobile devices. Therefore, existing technologies cannot provide a clear and operable risk measurement system, making it difficult to assist users in the field of intelligent mobile application development to effectively identify and address potential data risks. Summary of the Invention

[0003] Based on this, in view of the above technical problems, it is necessary to provide a risk assessment method for data dependency relationships in intelligent mobile application development that can assist users in the field of intelligent mobile application development to effectively identify and address potential data risks.

[0004] A risk assessment method for data dependency relationships in intelligent mobile application development, the method comprising: Obtain data from the HF platform to construct a dataset; preprocess the dataset to obtain a preprocessed dataset; define six entities according to the preprocessed dataset, including models, datasets, developers, annotators, tasks, and papers; construct a data dependency network based on the six entities and the key attributes between the entities; Build a visualization tool based on the Neo4j database to identify the structure and evolution information in the data dependency network through relationship queries, developer queries, and node path queries; the structure and evolution information includes vulnerable nodes and risk paths; Use the structure and evolution information to design quantitative indicators to obtain credibility risk indicators and vulnerability indicators; conduct a risk assessment of the data dependency relationship based on the credibility risk indicators and vulnerability indicators.

[0005] In the above method for risk assessment of data dependency relationships in intelligent mobile application development, this application can comprehensively analyze the dependency relationships between various components in the open-source model ecosystem by constructing a comprehensive data dependency network. By using the Neo4j database to manage complex dependency relationships and combining dynamic evolution analysis, it can accurately capture the changing trends of data and model dependency relationships. This method solves the problem that traditional methods cannot effectively handle large-scale and dynamically changing dependency relationships, greatly improving the ability to understand and monitor the structure of the dependency network. At the same time, "credibility risk" and "vulnerability" quantitative indicators are designed to make the risk assessment more accurate and intuitive. These two new indicators quantitatively evaluate the risks of data and models from multiple dimensions such as attribute integrity, timeliness, community feedback, and sensitivity of external dependencies, helping managers timely discover potential vulnerable nodes or untrusted dependencies, avoiding the defect that traditional methods cannot comprehensively evaluate risks, and thus enhancing the sustainability and stability of the open-source model ecosystem. Through the innovative combination of quantitative evaluation and dynamic analysis, the deficiencies in data dependency relationship management and risk assessment in the prior art are solved, not only improving the stability and reliability of the system, but also providing an effective technical means for the healthy management of the open-source model ecosystem in intelligent mobile application development, and being able to assist users in the intelligent mobile application development field to effectively identify and respond to potential data risks. Description of the Drawings

[0006] Figure 1 It is a schematic flowchart of a method for risk assessment of data dependency relationships in intelligent mobile application development in an embodiment; Figure 2 It is a schematic diagram of a data dependency network in an embodiment; Figure 3 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments

[0007] To make the objectives, technical solutions and advantages of this application more clear and understandable, the following further details this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not used to limit this application.

[0008] In one embodiment, as Figure 1 shown, a method for risk assessment of data dependency relationships in intelligent mobile application development is provided, including the following steps: Step 102, obtain the data construction dataset of the HF platform; preprocess the dataset to obtain the preprocessed dataset; define six entities according to the preprocessed dataset, including models, datasets, developers, annotators, tasks, and papers; construct a data dependency network based on the six entities and the key attributes between the entities.

[0009] By obtaining the data construction dataset of the HF (Hugging Face) platform, the data from the HF (Hugging Face) platform, on which more than 400,000 AI models and 100,000 datasets are hosted, with frequent data updates and rich resources. Data from the HF community is selected, and a data dependency network is constructed to comprehensively capture the dependency relationships and their evolution characteristics among various components in the open-source model ecosystem.

[0010] After preprocessing the dataset, six entities are defined, and a data dependency network is constructed based on these entities and their key attributes. Data dependency usually refers to models or datasets that declare some data information (such as the author of the model, the license of the model, etc., the author of the dataset, the license of the dataset, the dataset of the model, the base model of the model), etc. On the one hand, due to the dependency relationships, and on the other hand, due to the data, complex data dependency relationships are formed. This enables the various components and their dependency relationships in the originally complex and disorderly open-source model ecosystem to be presented in a structured manner, covering more comprehensively the associations in aspects such as data and models in intelligent mobile application development compared to traditional methods, thus effectively managing complex data dependency relationships. In the development of mobile image editing applications, the relationships between the model and the image dataset used for training, the improvements made by developers to the model, the annotations of data by annotators, etc., can all be clearly reflected in this data dependency network, facilitating developers to understand the dependency context throughout the development process.

[0011] Step 104, construct a visualization tool based on the Neo4j database to identify the structure and evolution information in the data dependency network through relationship queries, developer queries, and node path queries; the structure and evolution information include vulnerable nodes and risk paths.

[0012] Neo4j is selected as the underlying database management platform. Neo4j can effectively manage complex data relationships, which makes it very suitable for describing the dependency structure in the model ecosystem. The overall architecture of the data dependency network is as Figure 2 shown. When constructing the data dependency network, this application defines six main entities: "Model", "Dataset", "Author (i.e., developer)", "Annotator", "Task", and "Paper". Each entity has multiple key attributes. For example, the author, creation time, base model, and associated datasets of the model; the dataset contains information such as creation time, download times, annotator, and license. These attributes come from the tags on the HF platform and are mainly extracted from the model card and dataset card, forming the basis of the data dependency network. To deeply analyze the evolution of the data dependency network, the network is divided into multiple time windows, and based on the creation timestamps of the entities, how the dependencies change over time is observed. This process not only reveals the increasingly deepening dependency relationship between the model and the dataset, but also reflects the continuous reuse and interaction of the dataset and the model. By analyzing the network structure within different time windows, the dynamic evolution characteristics of the dependency relationship can be deeply understood.

[0013] According to the label description, the relationships between datasets are divided into two types: "SOURCE DATAFROM" and "EXTEND DATA FROM". The former points to the original dataset, while the latter indicates that the dataset is an extension or derivative based on other datasets. Through this label information, this application can capture the full picture of data dependencies, thus providing support for subsequent risk assessment.

[0014] Based on the Neo4j database, a visualization tool is constructed. Through methods such as relationship query, developer query, and node path query, the structure and evolution information in the data dependency network can be identified, including vulnerable nodes and risk paths. The Neo4j database is good at processing complex graph-structured data and can efficiently manage the many-to-many and dynamically changing dependency relationships that exist in large numbers in the open-source model ecosystem. The visualization tool allows developers to more intuitively see this information and timely discover potential problems in the dependency relationship.

[0015] During the frequent update and iteration of intelligent mobile applications, when the change in the user behavior data pattern leads to a change in the dependency relationship on the dataset, the affected nodes and paths can be quickly located through the visualization tool, so as to timely adjust the development strategy, solving the problem that traditional methods cannot respond to dynamic changes in a timely manner.

[0016] Step 106: Design quantification metrics using structural and evolutionary information to obtain credibility risk metrics and vulnerability metrics; conduct data dependency risk assessment based on the credibility risk metrics and vulnerability metrics.

[0017] The "credibility risk" and "vulnerability" quantification metrics designed in this application. These two metrics quantify and evaluate the risks of data and models from multiple dimensions, such as attribute integrity (e.g., model author information, license integrity), timeliness (the impact of the time after the model or dataset is released on its credibility and vulnerability), community feedback (the average number of "likes" per day reflects community recognition), and sensitivity of external dependencies (the impact of the node's dependency relationship with upstream nodes on vulnerability). It focuses on evaluating the subjective risks of the system (i.e., the autonomy and independence of the system when facing the external environment) and objective risks (i.e., the vulnerability of the system when affected by external changes).

[0018] Different from the traditional method that focuses on qualitative analysis and only pays attention to the compliance analysis of data sources and authorizations, these quantification metrics can extract quantifiable risk points from the full-process dependencies of the intelligent mobile application model construction, making the risk assessment more accurate and intuitive. When using the model for security risk assessment in mobile payment applications, the "credibility risk" metric can evaluate the credibility of the model itself, and the "vulnerability" metric can evaluate the stability of the model when external dependencies change, helping managers timely discover potential vulnerable nodes or untrusted dependencies, and effectively identify and respond to potential data risks.

[0019] In the above-mentioned data dependency relationship risk assessment method for intelligent mobile application development, this application can comprehensively analyze the dependency relationships among various components in the open-source model ecosystem by constructing a comprehensive data dependency network. By using the Neo4j database to manage complex dependency relationships and combining dynamic evolution analysis, it can accurately capture the changing trends of data and model dependency relationships. This approach solves the problem in traditional methods of being unable to effectively handle large-scale and dynamically changing dependency relationships, greatly improving the understanding and monitoring capabilities of the dependency network structure. At the same time, "credibility risk" and "vulnerability" quantification indicators are designed to make the risk assessment more accurate and intuitive. These two new indicators quantitatively evaluate the risks of data and models from multiple dimensions such as attribute integrity, timeliness, community feedback, and sensitivity of external dependencies, helping managers promptly discover potential vulnerable nodes or untrustworthy dependencies, avoiding the defect of traditional methods being unable to comprehensively evaluate risks, and thus enhancing the sustainability and stability of the open-source model ecosystem. Through the innovative combination of quantitative evaluation and dynamic analysis, the deficiencies in data dependency relationship management and risk assessment in the prior art are solved, not only improving the stability and reliability of the system, but also providing an effective technical means for the healthy management of the open-source model ecosystem in intelligent mobile application development, and being able to assist users in the field of intelligent mobile application development in effectively identifying and coping with potential data risks.

[0020] In one embodiment, the data set is preprocessed to obtain a preprocessed data set, including: After converting the CreatedAt field in the data set to the date and time format and extracting the developer information, data of the Space type and sparse or irrelevant fields, including the SHA hash value and redundant tags, are removed to obtain the preprocessed data set.

[0021] In a specific embodiment, during the data collection process, the HFCommunity data set is used, which covers three types of repositories on the HF platform: models, data sets, and spaces. To improve the accuracy of data analysis, the following steps are taken when processing the data: Convert the "CreatedAt" field to the date and time format; extract the developer information; and remove sparse or irrelevant fields such as the SHA hash value and redundant tags (such as "tag-TikTok" and "tag-Twitter").

[0022] Since repositories of the "Space" type mainly serve as examples of user interfaces, application displays, and model interactions and do not make a significant contribution to the analysis of data dependency relationships in the model ecosystem, these data are excluded.

[0023] In one embodiment, a quantification index is designed using structural and evolutionary information to obtain a credibility risk index and a vulnerability index, including: Based on the structural and evolutionary information, the credibility risk of data nodes is evaluated from three aspects: attribute integrity, timeliness, and the user perspective, and an exponential function and a logarithmic function are combined to construct a credibility risk index.

[0024] In a specific embodiment, for a given data node, its "credibility risk" is mainly evaluated from three aspects: Attribute integrity. Node attributes (such as author, license) and relationship attributes (such as source dataset or model) are crucial for ensuring legality and transparency. The absence of attributes such as license information may affect compliance. For almost isolated nodes (with only one inbound or outbound edge), intrinsic attributes are the main concern. For nodes in a dependency chain, the risk associated with incomplete attributes may spread through the entire chain, affecting overall reliability.

[0025] Timeliness. After a model or dataset is released, the corresponding real-world scenarios are often dynamic, which may lead to concept drift or data bias. Therefore, timeliness is crucial for the "credibility risk" of models and datasets.

[0026] From the user's perspective, nodes with a higher degree of internalization and more "likes" are usually subject to more frequent review and verification, especially within high-frequency communities. User feedback can improve the quality of the file and increase opportunities for improvement.

[0027] When constructing the "credibility risk" index, an exponential function and a logarithmic function are combined to effectively capture key risk factors and their interactions. The exponential function amplifies the impact of missing attributes, such as the absence of author information or license, because these absences seriously undermine the legality and credibility of the node. Therefore, a non-linear method is needed to accurately represent this potential threat. The logarithmic function is used to describe the gradually accumulating risks, such as time-related effects, so as to capture the gradual changes within the system and avoid a sudden drop in credibility due to information obsolescence. The average number of "likes" per day is also incorporated into the risk measure to comprehensively reflect the community's recognition. Relying solely on the number of "likes" may lead to biased evaluations, especially being disadvantageous to newer repositories due to time accumulation. By combining "likes" with "time", it is possible to better balance popularity and timeliness and reduce biases related to time differences. The integration of factors includes multiplication and division: multiplication captures the cumulative effects of different risk factors, while division helps control the overall risk level and prevent linear or excessive growth of risks. In addition, the number of subordinate nodes is added to the denominator to reflect the supervisory role of downstream nodes.

[0028] In one embodiment, the credibility risk indicator is: ; Wherein, represents the developer of the data node, and the developer is part of the node attributes, represents the license, and the license information is also one of the node attributes, represents the time interval, reflecting the time span from the release of the model or dataset to the current time, represents the number of "likes" per day on average, reflecting the community's recognition, represents the weight coefficient related to the node attributes, represents the time base value.

[0029] In one embodiment, according to the outbound relationship of the nodes in the structure and evolution information, that is, the dependency relationship of the nodes on the upstream nodes, and the credibility risk of the nodes themselves, the vulnerability indicator is obtained through relevant function mapping.

[0030] In a specific embodiment, the "vulnerability" indicator is different from traditional software. The situation in the model field is different: on the one hand, multiple dependency relationships can complement and enhance each other. For example, there may be many data sources with uneven quality, but this diversity helps to reduce the risk of bias caused by incomplete datasets. On the other hand, some data dependencies are one-time, such as datasets only used in the training phase. This means that "vulnerability" is not only the cumulative effect of upstream risks, but also closely related to the intrinsic quality of the nodes themselves.

[0031] To effectively address these unique challenges of artificial intelligence models, this application considers two key aspects.

[0032] (i) The outbound relationship of the nodes is considered, which represents the dependency relationship of the nodes on the upstream nodes. The more outbound links a node has, the more dependency relationships it maintains, which can increase or decrease its "vulnerability", depending on the complementarity of these dependency relationships. The importance of each relationship is further distinguished using the weighting factor W.

[0033] (ii) Different from traditional software measurement methods, this application emphasizes the credibility risk of the nodes themselves. Nodes with lower credibility risk indicate better data integrity, timeliness, and higher community recognition, so the "vulnerability" is lower. In addition, the "credibility risk" indicator is also used to balance the impact of upstream nodes on "vulnerability", and the internal quality of the nodes can offset part of the negative impact of external dependencies.

[0034] In one embodiment, the vulnerability indicator is: ; Wherein, Indicates the number of upstream nodes pointed to by the node, where the upstream nodes represent the relevant nodes in the outbound relationship. Indicates the number of upstream nodes pointed to by the node, where the upstream nodes represent the relevant nodes in the outbound relationship. Indicates the node The i th upstream 's credibility risk index value. Indicates the credibility risk index value of the current node 's credibility risk index value. Indicates the belief mapping function that affects the vulnerability of the credibility of the current node. Indicates the connection between the upstream node and the current node The weight of the relationship between them. Indicates the representative node The upstream node pointed to, that is, the node with an outbound relationship. Indicates the node for which the vulnerability index needs to be calculated currently.

[0035] In one embodiment, the relationship query is used to explore the graph structure of a specific type of relationship; the developer query is used to visualize a specific developer node and its relationships; the node path query is used to display the shortest path between nodes.

[0036] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,

[0037] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as shown in Figure 3As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it realizes a method for risk assessment of data dependency relationships in intelligent mobile application development. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, trackball, or touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0038] Those skilled in the art can understand that Figure 3 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0039] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0040] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0041] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A data dependency risk assessment method in intelligent mobile application development, characterized in that: The method comprises: Acquire data from the HF platform to build a data set; preprocess the data set to obtain a preprocessed data set; define six entities in the preprocessed data set, including models, data sets, developers, annotators, tasks, and papers; and build a data dependency network based on the six entities and key attributes between the entities; Building a visualization tool based on the Neo4j database to identify the structure and evolution information in the data dependency network through relationship query, developer query and node path query; the structure and evolution information includes vulnerable nodes and risk paths; Quantitative indicators are designed using the structure and evolution information to obtain credibility risk indicators and vulnerability indicators; and data dependency risk assessment is performed based on the credibility risk indicators and vulnerability indicators.

2. The method according to claim 1, characterized in that Preprocessing the data set to obtain a preprocessed data set includes: After converting the CreatedAt field in the dataset to date and time format and extracting the developer information, we remove Space type data and sparse or irrelevant fields, including SHA hash values ​​and redundant labels, to obtain the preprocessed dataset.

3. The method according to claim 1, characterized in that The structure and evolution information are used to design quantitative indicators to obtain credibility risk indicators and vulnerability indicators, including: According to the structural and evolutionary information, the credibility risk of data nodes is evaluated from three aspects: attribute integrity, timeliness and user perspective, and the credibility risk index is constructed by combining exponential function and logarithmic function.

4. The method according to claim 3, characterized in that The credibility risk indicators are: in, Indicates the developer of the data node. The developer is part of the node attributes. Represents a license. License information is also one of the node attributes. Indicates the time interval, reflecting the time span from the release of the model or dataset to the current time. Indicates the average number of "likes" per day, reflecting the community's recognition. Represents the weight coefficient related to the node attribute, Indicates the time base value.

5. The method according to claim 3, characterized in that: The method further comprises: According to the outbound relationship of nodes in the structure and evolution information, that is, the dependency of nodes on upstream nodes, and the credibility risk of the nodes themselves, the vulnerability index is obtained through correlation function mapping.

6. The method according to claim 5, characterized in that The vulnerability indicators are: in, Representation Node The number of upstream nodes pointed to, Representation Node The one pointed to Upstream The credibility risk index value of Indicates the current node The credibility risk index value of Represents the belief mapping function of how the credibility of the current node affects its vulnerability, Indicates connection to upstream node and the current node The weight of the relationship between Represents the node The upstream node pointed to is the node with an outbound relationship. Indicates the node for which the vulnerability index needs to be calculated.

7. The method according to claim 1, characterized in that The relationship query is used to explore the graph structure of a specific type of relationship; the developer query is used to visualize a specific developer node and its relationship; and the node path query is used to display the shortest path between nodes.

Citation Information

Patent Citations

  • Container cluster risk analysis and vulnerability assessment method and device based on environmental perception

    CN117874768A

  • Financial transaction anomaly detection and risk assessment method and device based on artificial intelligence

    CN119693111A

  • Multi-source software supply chain intelligent analysis method and system

    CN119720225A

  • Continuous vulnerability management for modern applications

    US20200242254A1

  • System for query-based interactive risk model analysis for secure software development

    US20220382878A1