Method of managing data, related apparatus and computer program product
By acquiring target data from multiple data sources, generating project ownership characteristics, and comparing similarities, the problem of data silos under traditional database management methods is solved, achieving efficient and high-quality unified data management and sharing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI HODE INFORMATION TECH CO LTD
- Filing Date
- 2026-02-14
- Publication Date
- 2026-06-09
Smart Images

Figure CN122173943A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, electronic device, computer-readable medium, and computer program product for managing data. Background Technology
[0002] With the rapid development of information technology and artificial intelligence, the workflows of various departments and processes in multiple fields such as the Internet, big data, artificial intelligence, and the Internet of Things have been digitized. This has led to an explosive growth in the amount of data output and processing, as more and more data needs to be collected, stored, analyzed, and transmitted.
[0003] Against this backdrop, while this explosive growth in data volume has brought opportunities for business development across various sectors, it has also presented significant challenges. For example, the increasing diversity and complexity of data types has rendered traditional database management methods and data storage architectures inadequate to meet the demands of data management and processing, making traditional data and database management approaches ineffective. Therefore, how to effectively manage and organize this data is a crucial and urgent issue that deserves attention. Summary of the Invention
[0004] This application provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for managing data, which can integrate data from different data sources based on the dimensions of the project, to eliminate data and information silos, achieve unified management and sharing of data, and thereby improve the efficiency and quality of data management and use.
[0005] One aspect of this application provides a method for managing data, comprising: obtaining target data corresponding to at least two data sources respectively; generating project affiliation features of the target data based on attribute information of the target data, wherein the attribute information includes at least one of the following: descriptive information corresponding to the data source, production time of the target data, and semantic information of the target data; comparing the similarity between the project affiliation features and project features of candidate projects to generate a similarity comparison result; determining the target project from the candidate projects based on the similarity comparison result; and storing the target data under the target project.
[0006] Another aspect of this application provides an apparatus for managing data, comprising: a data acquisition module configured to acquire target data corresponding to at least two data sources; an attribution feature generation module configured to generate item attribution features of the target data based on attribute information of the target data, wherein the attribute information includes at least one of the following: descriptive information corresponding to the data source, production time of the target data, and semantic information of the target data; a feature similarity comparison module configured to compare the similarity between the item attribution features and item features of candidate items, and generate a similarity comparison result; a target item determination module configured to determine a target item from the candidate items based on the similarity comparison result; and a data storage module configured to store the target data under the target item.
[0007] In another aspect of this application, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method for managing data as provided above.
[0008] Another aspect of this application provides a computer-readable storage medium having computer program instructions stored thereon, which can be executed by a processor to implement the method of managing data as provided above.
[0009] Another aspect of this application is a computer program product that includes a computer program having computer program instructions stored thereon, which, when executed by a processor, enables the implementation of the method for managing data as provided above.
[0010] The solution provided in this application involves obtaining target data from at least two data sources; generating project affiliation features of the target data based on the attribute information of the target data, wherein the attribute information includes at least one of the following: descriptive information corresponding to the data source, production time of the target data, and semantic information of the target data; comparing the similarity between the project affiliation features and the project features of candidate projects to generate a similarity comparison result; determining the target project from the candidate projects based on the similarity comparison result; and storing the target data under the target project. Thus, data from different data sources can be integrated based on the dimension of the project affiliation, eliminating data and information silos, achieving unified data management and sharing, and thereby improving the efficiency and quality of data management and use. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0013] Figure 1 A flowchart illustrating a data management process provided in one embodiment of this application;
[0014] Figure 2 A flowchart illustrating a process for acquiring target data, provided as an embodiment of this application;
[0015] Figure 3 A flowchart illustrating the process of managing data in a specific application scenario, as provided in another embodiment of this application;
[0016] Figure 4 This is a schematic diagram of a data management device provided in an embodiment of this application;
[0017] Figure 5 This is a schematic diagram of the structure of an electronic device suitable for implementing the solutions in the embodiments of this application.
[0018] The same or similar reference numerals in the accompanying drawings represent the same or similar parts. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] In a typical configuration of this application, the terminal and the service network devices each include one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0021] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0022] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer program instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, read-only optical disc (CD-ROM), digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0023] As discussed above, how to effectively manage and organize this data is a matter of concern and an urgent need.
[0024] In some solutions, it is possible to maintain a file server shared by multiple data sources, so that each data source can use the file server to store and share relevant data.
[0025] However, in this approach, data management and classification rely excessively on the storage actions of users (e.g., relevant personnel who provide data and use and manage data sources). This not only places high demands on users' operations but also incurs high operational costs. Furthermore, this approach often fails to achieve high-quality and efficient data management, especially in scenarios or departments where various data sources belong to different and widely different entities.
[0026] To address this issue, this application provides a method for managing data. This method obtains target data corresponding to at least two data sources; generates project affiliation features for the target data based on the attribute information of the target data, wherein the attribute information includes at least one of the following: descriptive information corresponding to the data source, production time of the target data, and semantic information of the target data; compares the similarity between the project affiliation features and project features of candidate projects to generate a similarity comparison result; determines the target project from the candidate projects based on the similarity comparison result; and stores the target data under the target project. Therefore, data from different data sources can be integrated based on the dimension of the project affiliation, eliminating data and information silos, achieving unified data management and sharing, and thus improving the efficiency and quality of data management and use.
[0027] In practical scenarios, the execution entity of this method can be a user device, a device composed of a user device and a network device integrated through a network, or an application running on the aforementioned devices. User devices include, but are not limited to, various terminal devices such as computers, mobile phones, tablets, smartwatches, and wristbands. Network devices include, but are not limited to, network hosts, single network servers, multiple network server sets, or cloud computing-based computer sets. Here, the cloud consists of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, consisting of a virtual computer composed of a group of loosely coupled computer sets.
[0028] When the executing entity is software, it can be installed in the electronic devices listed above. It can be implemented as multiple software programs or software modules, or as a single software program or software module, without specific limitations.
[0029] Figure 1 The present application illustrates a data management process 100, which includes at least the following processing steps:
[0030] (Step) S101: Obtain the target data corresponding to each of the at least two data sources;
[0031] In embodiments of this application, the execution entity (e.g., a server, such as a data management server) for performing the method or process of managing data can establish a data connection with the data sources in advance, or provide them with an interface capable of communicating and transmitting data. Accordingly, the execution entity can obtain the target data (i.e., the data that is expected to be stored in the execution entity, such as the server, so that it can be used and invoked by the data consumer) provided by at least two of these data sources.
[0032] For example, in an enterprise setting, the data source can be an internal data source within the enterprise. For instance, in an enterprise setting, the data source (or internal data source) can be based on data differences, specifically manifested as a requirement data source, a code data source, an evaluation data source, and so on.
[0033] Requirements data sources can be used, for example, by the requirements department within an enterprise to provide data such as requirements documents. Code data sources, on the other hand, can be used, for example, by the business departments or backend code development departments within an enterprise to provide data such as the code used to complete applications or products. Evaluation data sources can be used, for example, by the evaluation department within an enterprise to provide data such as evaluation results for code.
[0034] For example, in collaborative task scenarios, each data source can publish task data sources (e.g., provide or publish task requirements), process task data sources (actually provide task outputs), and review data sources (provide evaluation information on outputs), etc.
[0035] Accordingly, through this step, the executing entity can integrate the (target) data provided by each data source by communicating with at least two data sources, or in other words, multiple data sources, so that the target data and data sources can be connected, thereby solving the problem of data silos that may arise when each data source is independent and communicates with the executing entity separately.
[0036] It should be understood that the acquisition, storage, use, processing, transportation, provision and disclosure of any type of information, such as data, involved in the technical solutions of this application comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0037] S102, Based on the attribute information of the target data, generate the project affiliation characteristics of the target data;
[0038] In the embodiments of this application, after obtaining the target data based on the above S101, the executing entity can generate a corresponding project affiliation feature for each piece of target data based on its attribute information.
[0039] Attribute information can be information related to the source of the target data and the content it includes, used to represent, describe, indicate or characterize the target data, and can include at least one of the following: descriptive information corresponding to the data source, the production time of the target data, and semantic information of the target data.
[0040] For example, descriptive information can describe the name, port address, business department, etc. of the data source that provides the corresponding target data; production time can be the generation time of the corresponding target data (e.g., the moment when generation is completed), or the time period from the start of generation to the completion of generation; semantic information can be text content that is summarized based on the target data content to describe the content of the target data (e.g., describing the summarized requirements, the function of the code, the object being evaluated, etc.).
[0041] In some embodiments, based on the specific form of the target data and the differences in the content it includes, the attribute information may also include information that can reflect, describe, indicate, or characterize the target data, such as the data provider, data format, and purpose of data use.
[0042] In some embodiments, modeling can be performed in advance based on the structure, format, etc. of the target data that can be collected, so that after receiving the target data, the executing entity can determine the specific attribute information to be extracted based on the pre-modeling results. For example, the attribute information can be extracted from the description information of the corresponding data source, the production time of the target data, and the semantic information of the target data.
[0043] In this step, after extracting the attribute information, the executing entity can generate the project affiliation characteristics of the target data based on the attribute information.
[0044] Project attribution features can be a vector set, matrix, etc. Different dimensions within the vector set or matrix can be used to correspond to specific attribute information, allowing for a unified representation and aggregation of the target data's attribute information across different dimensions. For example, when simultaneously using the descriptive information of the corresponding data source, the production time of the target data, and the semantic information of the target data, the project attribution feature can include three dimensions. These three dimensions can then be used to record the corresponding data source description, production time, and semantic information of the target data, respectively.
[0045] It should be understood that the dimensional division of the project attribution feature usually corresponds one-to-one with the project features discussed below, so that in the next (step) S103, the project to which the target data belongs can be determined by comparing the similarity between the project attribution feature and the project features of the candidate projects.
[0046] Next, we will explain the actions implemented in S103 in a unified manner.
[0047] S103, compare the similarity between the project attribution characteristics and the project characteristics of the candidate projects, and generate similarity comparison results;
[0048] In the embodiments of this application, as discussed above, the executing entity can compare the project attribution features generated in S102 with the project features of the candidate projects to generate a similarity comparison result between the two.
[0049] Candidate projects are typically pre-categorized based on business and task requirements. For example, in an enterprise setting, candidate projects can be pre-categorized based on the company's internal situation (e.g., business situation). This approach allows for the summarization and unified management of data belonging to or related to the same project, using "project" as a metric or classification method to conveniently and uniformly access related data. This avoids data silos caused by data fragmentation within the same project, addressing the issue of relevant personnel being unable to easily access related data. For instance, in an enterprise setting, this method can connect users from different departments through "projects," enabling them to share and access project-related data in a timely manner.
[0050] In practice, projects can be further divided based on business departments and business situations to ensure that different product lines and business lines have independent data usage spaces, thereby enhancing the domain specificity of the data. For example, projects can be divided based on product lines such as video, live streaming, and literary creation.
[0051] After identifying the "project", the project characteristics can be determined based on the performance of the project in the "dimensions" involved in the aforementioned attribute information. For example, for project A, different dimensions can be used in its project characteristics to indicate the data sources, personnel, production time, etc. involved in the project. This allows the executing entity to organize and classify the data generated by these data sources and personnel during the production time into the project, so that the data associated with the project can be organized and aggregated across data sources and specific content forms (or formats) with "project" as the dimension.
[0052] Correspondingly, these "known" items can be referred to as candidate items (typically, the name of the candidate item can be based on user input or determined by parsing the content in the requirements document). Then, the implementing entity can generate a similarity comparison result by comparing the similarity between the item attribution characteristics and the item characteristics of the candidate items, and use the similarity comparison result to determine whether the currently acquired target data can be attributed to or correspond to a candidate item.
[0053] In some embodiments, during the determination of project features, they may be dimensions that are exactly the same as "project attribution features," or they may only refer to or include some dimensions of "project attribution features." Accordingly, if project features only refer to or include some dimensions of "project attribution features," the implementing entity can choose to perform similarity comparisons only on these included dimensions to determine the similarity comparison results. This avoids undesirable significant changes in similarity comparison results due to variations in the dimensions of project attribution features caused by data differences (e.g., temporary appearance or lack of some dimensions weakly related to "project"). It also allows for targeted and differentiated setting of corresponding project features for different projects, based on their respective dimensions of interest, rather than uniformly setting project features that include all dimensions.
[0054] S104, Based on the similarity comparison results, the target project is determined from the candidate projects;
[0055] In the embodiments of this application, after obtaining the similarity comparison result based on the above S103, the executing entity can use the similarity comparison result to determine the target project from the candidate projects. For example, the executing entity can use the candidate project with the highest similarity as the target project to be used to assign the target data.
[0056] In some embodiments, to avoid incorrect attribution due to the failure of "projects" (i.e., candidate projects) to be configured or synchronized in a timely manner, the executing entity may further determine in this step whether the highest similarity is greater than or equal to a similarity threshold (which can be set based on the standard that the target data can be reliably considered to belong to a project).
[0057] Accordingly, in this case, the executing entity can, as an alternative or additional step, first select whether the detection similarity comparison result indicates the existence of a target candidate item, the target similarity of which is greater than or equal to the similarity threshold and greater than other candidate items (i.e., its corresponding target similarity is the highest and is already greater than or equal to the similarity threshold).
[0058] If it exists, the executing entity can respond by setting the target candidate as the target project.
[0059] In some alternative ways of this embodiment, if the similarity comparison result indicates that there is no target candidate item, the executing entity may respond to this by choosing to generate the target item based on the attribute information (i.e., detaching from the existing candidate items and generating a new "item").
[0060] For example, the executing entity can choose to use the attribute information corresponding to the target data as the project features of the new target project. Alternatively, if there are multiple target data, the executing entity can first identify the target data that cannot be classified into the current candidate project, and then use the attribute information of the clustered target data, such as the average value of the attribute information in each dimension, to similarly generate the project features of the new target project.
[0061] This allows target data that cannot be reliably attributed to be managed by creating or adding new projects. This enables subsequent users to directly use this data by, for example, confirming new projects or partially changing record information. This allows the implementing entity to not only have the ability to explore potential projects that have not been pre-divided or identified, but also to have the corresponding management ability for the data under these potential projects, thus enabling better data management.
[0062] In some embodiments, when the implementing entity discovers and adds new projects, it can first name them according to a pre-determined temporary naming rule so that they can be perceived and identified by the user.
[0063] In some embodiments, the executing entity may also provide certain prompts based on the specific content included in the attribute information (or, in other words, the attribute information used to generate project characteristics), so that users can use the prompts to understand the details of the newly added project and avoid user confusion. Accordingly, after understanding the details of the newly added project based on the prompts, users can also adjust the recorded information, such as the project name and project characteristics, through interaction with the executing entity to meet their usage needs.
[0064] S105, Store target data under the target project.
[0065] In the embodiments of this application, after the target data is assigned to the target project in S104, the executing entity can assign the target data to the target project. For example, the executing entity can store the target data to the storage address corresponding to the target project, so as to store and manage the target data at the dimension of "project".
[0066] In practice, for a "project", multiple storage addresses can be pre-classified based on the data types it can include, so that the executing entity can adaptably classify and store the target data under the target project according to the data types of the target data.
[0067] In some embodiments, for these projects as a whole, the executing entity can establish the overall relationship between projects and target data in the form of a "graph structure." By adding new target data and establishing relationships between target data and projects, the executing entity can store and maintain target data at the "project" level. Similarly, the executing entity can subsequently maintain a "database" for storing and managing target data by continuously modifying and updating the graph structure. Thus, through this "graph structure" approach, the executing entity can manage target data through a graph-level incremental update mechanism, avoiding full reconstruction, maintaining real-time knowledge evolution, and improving data management quality.
[0068] In some embodiments, to improve the visualization of "data" for users and facilitate their reading and use, the executing entity can also choose to construct data belonging to the same project using terms. That is, "project" can be used as a term, and data associated with that project can be displayed as term information under that project. For example, the executing entity can also set corresponding information fields for data types, so that it can maintain and construct project-based data sets by inserting term information.
[0069] In this scenario, the executing entity can first select and retrieve the terminology information corresponding to the target item; then, based on the data type of the target data, determine the target information field within the terminology information; finally, store the target data in the target information field. Thus, by utilizing terms to provide data maintenance and association, users can efficiently and cost-effectively obtain the data they need through terms.
[0070] In some embodiments, to avoid terms being unavailable or occupied for a long time due to updating the terms, the executing entity may choose to first compare the content difference between the existing data in the target information field and the target data during the process of storing target data in the target information field, and generate a content difference comparison result.
[0071] For example, the implementing entity can determine the "content difference" by the semantic similarity between the target data and the existing data, or by the degree of support that the target data can be proven and supported by the existing data. For example, the result of 1-semantic similarity, 1-degree of proof, and support can be used as the "content difference".
[0072] Then, if the content difference comparison result indicates that the content difference between the two is greater than or equal to the difference threshold (usually, this can be preset based on the standard that the two have a large difference and need to be updated immediately), the executing entity can respond by storing the target data in the target information field, that is, storing the target data "immediately". This ensures that target data with larger differences, which are more likely to provide more and newer data content, can be maintained and stored in a timely manner.
[0073] In such cases, in some embodiments, if the content difference comparison result indicates that the content difference between the two is less than the aforementioned difference threshold, the executing entity may also respond by storing the target data in the data cache.
[0074] This cache can be used to write the stored target data to the corresponding target item when the amount of stored target data is greater than or equal to a storage threshold (for example, it can be set based on the standard that the accumulated data amount is sufficient to trigger a write or storage action). Thus, this cache can be used to accumulate target data that has already been largely recorded and whose update value may be low, and then trigger storage and update actions only when the data amount is sufficient, avoiding frequent updates to entries that could lead to entries being occupied and unusable for a long time.
[0075] In some embodiments, during the process of writing and storing target data, the executing entity may also choose to use idempotent updates or incremental updates to balance the needs of update resource usage and data consistency.
[0076] In some embodiments, the executing entity may also choose to select data that needs to be stored and written immediately based on data type, as well as data that can be stored and written in "batch" using the cache, in order to balance the relationship between data timeliness and term usage quality by combining the differences between data types.
[0077] Subsequently, the data management method provided in this application obtains target data corresponding to at least two data sources; generates project affiliation features of the target data based on the attribute information of the target data, wherein the attribute information includes at least one of the following: descriptive information corresponding to the data source, production time of the target data, and semantic information of the target data; compares the similarity between the project affiliation features and the project features of candidate projects to generate a similarity comparison result; determines the target project from the candidate projects based on the similarity comparison result; and stores the target data under the target project. Thus, it is possible to integrate data from different data sources based on the dimension of the project affiliation, thereby eliminating data and information silos, achieving unified data management and sharing, and improving the efficiency and quality of data management and use.
[0078] In some embodiments, as discussed above, due to differences between data sources (e.g., differences between departments or tasks), the content, form, and structure of the data provided by the data sources often vary significantly. Therefore, in order to reduce the difficulty of data synchronization from the data sources and improve data management and execution, the executing entity may choose to treat the directly obtained data as initial data during the process of obtaining target data from the data source, and then actually obtain the target data by processing the initial data, such as filtering or supplementing it.
[0079] Accordingly, in some embodiments, when the executing entity obtains the target data corresponding to the data source from at least two data sources, it may first select to obtain the initial data corresponding to the data source from at least two data sources.
[0080] Then, based on the data type of the initial data, the data filtering rules corresponding to the initial data are determined.
[0081] Specifically, in this step, the data filtering strategies to be executed can be determined in advance for different data types, such as noise reduction, template removal, paragraph splitting, outline extraction, entity recognition, field mapping, etc. For the required text, the corresponding data filtering strategies can be outline extraction, entity recognition, etc.
[0082] Then, the executing entity can filter target data from the initial data based on data filtering rules. Thus, by using data filtering strategies corresponding to the data type, data filtering operations such as noise reduction and content simplification can be performed in a compliant manner to improve the data quality of the stored target data.
[0083] In some embodiments, different data filtering strategies can be concretely trained as corresponding filtering models or agents, so that subsequent data filtering can be efficiently and effectively completed by calling the corresponding filtering model or agent to apply the data filtering strategy. An intelligent agent is a system or entity capable of perceiving its environment and autonomously making decisions and executing tasks. It can make judgments and choices based on its own goals and changes in the external environment, and take actions to achieve a specific purpose. For example, an intelligent agent used to complete a data filtering strategy can be pre-trained with unprocessed initial sample data as input and processed target sample data as output, enabling it to obtain target sample data by deleting unwanted portions of the initial sample data.
[0084] In some embodiments, in addition to data filtering, the executing entity can also manage "content completeness" to avoid the influx of low-quality initial data with low content completeness, which would result in a waste of computing resources.
[0085] To better understand the complete process of acquiring target data in this situation, please refer to [the relevant documentation / reference]. Figure 2 . Figure 2 A flowchart of a process 200 for acquiring target data according to an embodiment of this application is shown. The process 200 can be used as an alternative or alternative implementation of the above-described S101.
[0086] Process 200 may specifically include the following steps:
[0087] S201, Obtain the initial data corresponding to each of the at least two data sources;
[0088] As discussed above, in this step, the executing entity can first use data obtained from at least two data sources as initial data, rather than target data.
[0089] S202, based on the data type of the initial data, call the content completeness evaluation model corresponding to the initial data;
[0090] Specifically, similar to the discussion above, corresponding content completeness evaluation models can be pre-configured based on data types. For example, for requirement documents, a content completeness evaluation model can be configured to determine and generate content completeness from semantic information and semantic completeness; similarly, for "code," a content completeness evaluation model can be configured to determine and generate content completeness from functional completeness.
[0091] Accordingly, in this step, the executing entity can call the content completeness evaluation model corresponding to the initial data based on the data type of the initial data, and then process the initial data using the content completeness evaluation model by executing the following S203 to generate the corresponding completeness evaluation result. This completeness evaluation result can usually be a specific "completeness value".
[0092] S203, use the content completeness evaluation model to process the initial data and generate the corresponding completeness evaluation results;
[0093] Next, if the completeness evaluation result indicates a completeness greater than or equal to the completeness threshold (which can usually be set based on the standard that the completeness of the initial data is considered to meet the usage requirements), the executing entity can respond to this by selecting to execute S204 and performing the data filtering process as discussed above (e.g., "determine the data filtering rules corresponding to the initial data based on the data type of the initial data").
[0094] S204, Determine the data filtering rules corresponding to the initial data.
[0095] This allows the implementing entity to remove incomplete or low-quality data inflows after obtaining the initial data through a completeness check, thus avoiding the waste of computing resources and preventing data pollution caused by these low-quality data inflows.
[0096] In some optional implementations of this embodiment, if the initial data is not complete enough, the executing entity may also choose to generate supplementary prompts based on content types that can improve completeness (e.g., content that can supplement semantic information, such as explanatory information for a certain "entity"; or, for example, missing functions in the code). These supplementary prompts can be used to assist the user in making targeted adjustments and supplements to the missing content, enabling the user to complete the data provision work more efficiently and with higher quality.
[0097] Accordingly, process 200 may also include steps S205 and S206. Step S205 may be performed by the executing entity after step S203 if the completeness indicated by the completeness evaluation result is less than the completeness threshold.
[0098] S205, Based on the content type that can improve the completeness, generate supplementary prompt information;
[0099] Specifically, as discussed above, the implementing entity determines the types of content that can improve completeness (or the types of content missing in the current initial data) by obtaining the model parameters of the content completeness evaluation model in generating the completeness evaluation results, or by requesting the "content completeness evaluation model" to provide the content it believes is missing (for example, when the content completeness evaluation model is built on a large language model).
[0100] Then, the executing entity can generate supplementary prompts based on the content type to indicate that the content type. For example, it could be in the form of "The content is incomplete because the provided data lacks explanation of 'entity XX', so it is recommended to supplement the relevant content".
[0101] S206, return supplementary prompts and initial data to the target device associated with the data source that provided the initial data.
[0102] Specifically, after generating supplementary prompt information based on the above S205, the executing entity can return the supplementary prompt information and the corresponding initial data to the target device associated with the data source that provided the initial data (for example, in an enterprise scenario, it could be a terminal device used by users within the enterprise that can use it as a data source or communicate with the data source to provide initial data), so that the data source can understand the need to improve the data in a timely and cost-effective manner.
[0103] Building upon any of the above embodiments, to further enhance the usability of the stored target data, the executing entity can also obtain a list of supported models associated with the target project after storage is complete, such as a Large Language Model (LLM). The model list records the target models associated with the target project, as well as the target models' format requirements for model output (e.g., vectors, data embeddings, matrices, etc.).
[0104] For example, the model list can specifically record that the target models required for project A are Model A, Model B, and Model C, and that Model A requires data XX for project A, with the required data format being "vector form" and "data embedding".
[0105] In this case, during or after storing the target data, the executing entity can process the target data into the required format or form (e.g., the "vector form" or "data embedding" mentioned above) based on the information recorded in the model support list. That is, the target data is processed into indexed information corresponding to the target format requirements.
[0106] Then, after processing the target data and converting it into indexed information, the executing entity can create model index information (e.g., "Query") corresponding to the indexed information. This allows the corresponding model to directly obtain and index the target data in the required data format and form through the model index information. For example, a database can be constructed to implement Retrieval-augmented Generation (RAG).
[0107] This allows subsequent models to directly obtain the required data format and form when utilizing the stored target data, and avoids erroneous searches and matches such as synonyms due to unclear index pointing, thereby improving the performance of the execution entity in providing target data.
[0108] It should be understood that such index correspondence can also be one-to-many. For example, based on the semantic relevance of the indexed information and the Top-k retrieval method, the "k" indexed results can be associated with a model index, so that the model index can comprehensively utilize multiple data, or in other words, more effectively complete the model task by aggregating and jointly utilizing multiple data.
[0109] To enhance understanding, this application also provides a specific implementation scheme based on a particular application scenario. Please refer to it. Figure 3 , Figure 3 This is a flowchart of a data management process 300 implemented in a specific application scenario, as provided in an embodiment of this application.
[0110] In process 300, server 310 may be used as the execution entity of the process for managing data, for obtaining data from data sources (e.g., terminal devices 311, 312 and 313) and managing them.
[0111] Specifically, in process 300, server 310 can first obtain initial data from terminal devices 311, 312 and 313, which are data sources, by executing S301. For example, it can obtain initial data 321 from terminal device 311, initial data 322 from terminal device 312 and initial data 323 from terminal device 313.
[0112] It should be understood that the number of “data sources” and the number of initial data mentioned above are merely illustrative examples for ease of understanding and are not intended to limit the specific value of the “number”.
[0113] Then, server 310 can execute S302 to determine the data filtering rules corresponding to each initial data based on the data type of the (initial) data. For example, based on the data type of initial data 321, the corresponding data filtering rule 331 is determined; based on the data type of initial data 322, the corresponding data filtering rule 332 is determined; and based on the data type of initial data 323, the corresponding data filtering rule 333 is determined.
[0114] Next, server 310 can execute S303 to filter target data from initial data based on data filtering rules. For example, different data filtering rules can be used to execute data filtering rules such as noise reduction, deduplication, entity extraction, etc. For example, data filtering rule 331 can be used to process and extract initial data 321 into target data 341, data filtering rule 332 can be used to process and extract initial data 322 into target data 342, and data filtering rule 333 can be used to process and extract initial data 323 into target data 343.
[0115] Next, as discussed above, server 310 can continue to execute S304 to generate project ownership features of the target data based on the attribute information of the target data. For example, target data 341 can be generated with project ownership feature 351, target data 342 can be generated with project ownership feature 352, and target data 343 can be generated with project ownership feature 353.
[0116] Then, server 310 can execute S305 to compare item attribution features 351, 352, and 353 with item features 361, 362…36N (where N is a positive integer) to generate similarity comparison results. In this process, item features 361, 362…36N can correspond to pre-determined "items," or "candidate items," so that the similarity comparison results can be used to determine whether target data 341, 342, and 343 can be classified into these items.
[0117] For the purpose of keeping the content of the diagram concise, Figure 3 The example only uses target data 341. For instance, after server 310 executes S305, for target data 341, the comparison result of its corresponding project attribution feature 351 may show that project attribution feature 351 has the highest (target) similarity to project feature 361, and this (target) similarity is greater than or equal to a similarity threshold. Accordingly, based on this comparison result, target data 341 can be determined to belong to the "project" corresponding to project feature 361, for example, project 371.
[0118] Finally, based on this attribution result, server 310 can choose to execute S307 to store target data 341 under project 371. Similarly, for other target data (e.g., target data 342 and 343), server 310 can also determine the project to which it belongs (or create a new project for them) in a similar manner based on the comparison results of their corresponding project attribution features 352 and 353, which will not be elaborated here.
[0119] This application also provides an apparatus for managing data, the structure of which is as follows: Figure 4 The apparatus 400 shown includes: a data acquisition module 410 configured to acquire target data corresponding to at least two data sources; an attribution feature generation module 420 configured to generate item attribution features of the target data based on attribute information of the target data, wherein the attribute information includes at least one of the following: descriptive information corresponding to the data source, production time of the target data, and semantic information of the target data; a feature similarity comparison module 430 configured to compare the similarity between the item attribution features and the item features of candidate items, and generate a similarity comparison result; a target item determination module 440 configured to determine the target item from the candidate items based on the similarity comparison result; and a data storage module 450 configured to store the target data under the target item.
[0120] This embodiment exists as a device embodiment corresponding to the above method embodiment. The device for managing data provided in this embodiment can integrate data from different data sources based on the dimension of the project, so as to eliminate data and information silos, realize unified management and sharing of data, and thereby improve the efficiency and quality of data management and use.
[0121] In some embodiments, the data acquisition module 410 includes: an initial data acquisition submodule configured to acquire initial data corresponding to at least two data sources respectively; a filtering rule determination submodule configured to determine a data filtering rule corresponding to the initial data based on the data type of the initial data; and a data filtering submodule configured to filter target data from the initial data based on the data filtering rule.
[0122] In some embodiments, the filtering rule determination submodule includes: an evaluation model invocation unit, configured to invoke a content completeness evaluation model corresponding to the initial data based on the data type of the initial data; a completeness evaluation unit, configured to process the initial data using the content completeness evaluation model to generate a corresponding completeness evaluation result; and a filtering rule determination unit, configured to determine a data filtering rule corresponding to the initial data in response to the completeness evaluation result indicating that the completeness is greater than or equal to the completeness threshold.
[0123] In some embodiments, the filtering rule determination submodule may further include: a supplementary prompt generation unit, configured to generate supplementary prompt information based on content types that can improve completeness in response to a completeness evaluation result indicating that the completeness is less than a completeness threshold; and a prompt and initial data return unit, configured to return the supplementary prompt information and initial data to a target device associated with the data source providing the initial data.
[0124] In some embodiments, the target item determination module 440 is further configured to, in response to a similarity comparison result indicating the existence of a target candidate item, use the target candidate item as the target item, wherein the target similarity corresponding to the target candidate item is greater than or equal to a similarity threshold and is greater than that of other candidate items.
[0125] In some embodiments, the apparatus 400 further includes a target item generation module, configured to generate a target item based on attribute information in response to a similarity comparison result indicating that no target candidate item exists.
[0126] In some embodiments, the data storage module 450 includes: a term information retrieval submodule, configured to retrieve term information corresponding to the target item; an information field determination submodule, configured to determine the target information field based on the data type of the target data in the term information; and a data storage submodule, configured to store the target data in the target information field.
[0127] In some embodiments, the data storage submodule includes: a difference comparison unit configured to compare the content difference between existing data in the target information field and target data, and generate a content difference comparison result; and a first data storage unit configured to store the target data in the target information field in response to the content difference comparison result indicating that the content difference between the two is greater than or equal to a difference degree threshold.
[0128] In some embodiments, the data storage submodule may further include: a second data storage unit configured to store the target data in a data cache in response to a content difference comparison result indicating that the content difference between the two is less than a difference threshold, wherein the cache is used to write the stored target data to the corresponding target item when the amount of stored target data is greater than or equal to the storage threshold.
[0129] In some embodiments, the apparatus 400 further includes: a support list acquisition module configured to acquire a model support list associated with a target project, wherein the model list records the target models associated with the target project and the format requirements of the target models for model output; an indexed information processing module configured to use the model support list to process the target data into indexed information corresponding to the format requirements of the target format; and an indexed information creation module configured to create model indexed information corresponding to the indexed information.
[0130] Based on the same concept, this application also provides an electronic device, a readable storage medium, and a computer program product. The method corresponding to the electronic device can be the data management method in the foregoing embodiments, and its problem-solving principle is similar to that method. The electronic device provided in this application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the methods and / or technical solutions of the foregoing embodiments of this application.
[0131] Electronic devices can be user devices, or devices composed of user devices and network devices integrated through a network, or applications running on the aforementioned devices. User devices include, but are not limited to, various terminal devices such as computers, mobile phones, tablets, smartwatches, and wristbands. Network devices include, but are not limited to, network hosts, single network servers, multiple network server sets, or cloud computing-based computer sets, and can be used to implement some processing functions when setting an alarm clock. Here, the cloud consists of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, consisting of a virtual computer composed of a group of loosely coupled computer sets.
[0132] Figure 5 The diagram illustrates the structure of an electronic device suitable for implementing the methods and / or technical solutions in the embodiments of this application. The electronic device 500 includes a Central Processing Unit (CPU) 501, which can perform various appropriate actions and processes based on a program stored in a Read Only Memory (ROM) 502 or a program loaded from a storage portion 508 into a Random Access Memory (RAM) 503. The RAM 503 also stores various programs and data required for system operation. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus 504. An Input / Output (I / O) interface 505 is also connected to the bus 504.
[0133] The following components are connected to I / O interface 505: an input section 506 including a keyboard, mouse, touchscreen, microphone, infrared sensor, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), LED display, OLED display, etc., and speakers, etc.; a storage section 508 including one or more computer-readable media such as hard disk, optical disk, magnetic disk, semiconductor memory, etc.; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet.
[0134] In particular, the methods and / or embodiments in this application can be implemented as computer software programs. For example, the embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 501, it performs the functions defined in the methods of this application.
[0135] Another embodiment of this application provides a computer-readable storage medium and a computer program product having computer program instructions stored thereon, which can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of this application described above.
[0136] Specifically, this embodiment may employ any combination of one or more computer-readable media. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, a system, apparatus, or device that is, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0137] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0138] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0139] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—as well as conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0140] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0141] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0142] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules and units is only a logical functional division, and in actual implementation, there may be other division methods. Taking units as examples, multiple units or page components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0143] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0144] Furthermore, the functional modules and units in the various embodiments of this application can be integrated into one processing module or unit, or each module or unit can exist physically separately, or two or more units can be integrated into one module or unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules and units.
[0145] The integrated modules and units implemented as software functional modules and units described above can be stored in a computer-readable storage medium. These software functional modules and units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
[0147] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device through software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any specific order.
Claims
1. A method for managing data, characterized in that, include: Obtain the target data corresponding to each of at least two data sources; Based on the attribute information of the target data, a project affiliation feature of the target data is generated, wherein the attribute information includes at least one of the following: descriptive information corresponding to the data source, the production time of the target data, and the semantic information of the target data; Compare the similarity between the project attribution characteristics and the project characteristics of the candidate projects to generate similarity comparison results; Based on the similarity comparison results, the target project is determined from the candidate projects; The target data is stored under the target project.
2. The method according to claim 1, characterized in that, The step of obtaining the target data corresponding to each of at least two data sources includes: Obtain the initial data corresponding to each of at least two data sources; Based on the data type of the initial data, determine the data filtering rules corresponding to the initial data; Based on the data filtering rules, target data is filtered out from the initial data.
3. The method according to claim 2, characterized in that, The step of determining the data filtering rules corresponding to the initial data based on the data type of the initial data includes: Based on the data type of the initial data, the content completeness evaluation model corresponding to the initial data is invoked; The initial data is processed using the content completeness evaluation model to generate corresponding completeness evaluation results; In response to the completeness evaluation result indicating that the completeness is greater than or equal to the completeness threshold, a data filtering rule corresponding to the initial data is determined.
4. The method according to claim 3, characterized in that, The method further includes: In response to the completeness evaluation result indicating that the completeness is less than the completeness threshold, supplementary prompts are generated based on content types that can improve the completeness. The supplementary prompt information and the initial data are returned to the target device associated with the data source that provided the initial data.
5. The method according to claim 1, characterized in that, The process of determining the target project from the candidate projects based on the similarity comparison results includes: In response to the similarity comparison result indicating the existence of a target candidate item, the target candidate item is selected as the target item, wherein the target similarity corresponding to the target candidate item is greater than or equal to a similarity threshold and is greater than that of the other candidate items.
6. The method according to claim 5, characterized in that, The method further includes: In response to the similarity comparison result indicating that the target candidate item does not exist, the target item is generated based on the attribute information.
7. The method according to claim 1, characterized in that, The storage of the target data under the target project includes: Retrieve the terminology information corresponding to the target item; The target information field is determined based on the data type of the target data in the term information; The target data is stored in the target information field.
8. The method according to claim 7, characterized in that, The step of storing the target data in the target information field includes: Compare the existing data in the target information field with the target data to generate a content difference comparison result; In response to the content difference comparison result indicating that the content difference between the two is greater than or equal to the difference degree threshold, the target data is stored in the target information field.
9. The method according to claim 8, characterized in that, The method further includes: In response to the content difference comparison result indicating that the content difference between the two is less than the difference threshold, the target data is stored in a data cache area, wherein the cache area is used to write the stored target data into the corresponding target item when the amount of stored target data is greater than or equal to the storage threshold.
10. The method according to any one of claims 1-9, characterized in that, The method further includes: Obtain a list of supported models associated with the target project, wherein the list of models records the target models associated with the target project and the format requirements of the target models for model output; Using the model support list, the target data is processed into indexed information corresponding to the target format required by the format; Establish model index information corresponding to the indexed information.
11. A device for managing data, characterized in that, include: The data acquisition module is configured to acquire target data corresponding to at least two data sources, respectively. The attribution feature generation module is configured to generate project attribution features of the target data based on the attribute information of the target data, wherein the attribute information includes at least one of the following: descriptive information corresponding to the data source, the production time of the target data, and the semantic information of the target data; The feature similarity comparison module is configured to compare the similarity between the project's attribution feature and the project features of the candidate projects, and generate a similarity comparison result. The target project determination module is configured to determine the target project from the candidate projects based on the similarity comparison results; The data storage module is configured to store the target data under the target project.
12. An electronic device, the electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 10.
13. A computer-readable medium having stored thereon computer program instructions that can be executed by a processor to implement the method as claimed in any one of claims 1 to 10.
14. A computer program product comprising a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 10.