Data processing method and device, electronic equipment and readable storage medium
By generating a local knowledge index on the terminal and performing anonymization processing, and uploading metadata to the server to build a global data graph, the problem of low information retrieval efficiency in multiple terminal devices is solved, and secure and efficient cross-device data management is achieved.
Patent Information
- Application Number
- CN202511712125.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-02-17
AI Technical Summary
In multi-device scenarios, users need to search for information from multiple devices, which is inefficient, and existing technologies cannot effectively protect user privacy and data security.
By generating a local knowledge index on the terminal, de-identifying it, and then uploading the metadata to the server, the server builds a global data graph, enabling secure and efficient information retrieval across devices.
It improves the efficiency of information retrieval in multi-terminal device scenarios, reduces the risk of data leakage, and protects user privacy.
Smart Images

Figure CN121542368A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of electronic equipment technology, specifically relating to a data processing method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] With the rapid development of mobile internet and IoT technologies, the types and number of smart terminal devices owned by individual users have increased significantly, resulting in an exponential increase in the amount of data generated. Currently, user data is no longer centrally stored on a single terminal device, but is widely distributed across multiple terminal devices such as smartphones, personal computers, tablets, and smart wearable devices, thus forming multiple physically isolated "data silos." In this context, when users need to find certain information, they must search sequentially across multiple terminal devices, leading to low information retrieval efficiency. Summary of the Invention
[0003] The purpose of this application is to provide a data processing method, apparatus, electronic device, and readable storage medium that can improve the efficiency of information retrieval in multi-terminal device scenarios.
[0004] In a first aspect, embodiments of this application provide a data processing method applied to a first terminal, the method comprising: Generate a local knowledge index based on local data; The local knowledge index is anonymized to obtain anonymized metadata; wherein the metadata includes descriptive information and location information; The metadata is uploaded to a server; wherein the server constructs a global data graph based on metadata from at least two terminals.
[0005] Secondly, embodiments of this application provide a data processing apparatus applied to a first terminal, the apparatus comprising: The generation module is used to generate a local knowledge index based on local data. The processing module is used to perform de-identification processing on the local knowledge index to obtain de-identified metadata; wherein, the metadata includes description information and location information; A sending module is used to upload the metadata to a server; wherein the server constructs a global data graph based on metadata from at least two terminals.
[0006] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the data processing method as described in the first aspect.
[0007] Fourthly, embodiments of this application provide a computer-readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the data processing method as described in the first aspect.
[0008] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the data processing method as described in the first aspect.
[0009] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the data processing method as described in the first aspect.
[0010] In this embodiment, the first terminal generates a local knowledge index based on local data; the local knowledge index is anonymized to obtain anonymized metadata, wherein the metadata includes descriptive information and location information; the metadata is uploaded to the server; and the server constructs a global data graph based on the metadata from at least two terminals. On one hand, the anonymization of the local knowledge index by the terminal effectively removes sensitive information that may leak user privacy, providing security protection before the data is uploaded to the server. This ensures that even if the metadata is illegally obtained during transmission or storage, attackers cannot obtain the user's sensitive information, reducing the risk of data leakage. On the other hand, users do not need to manage data on different terminals separately; they can search the entire data set simply through the global data graph on the server, improving information retrieval efficiency in multi-terminal device scenarios. Attached Figure Description
[0011] Figure 1 This is one of the flowcharts of a data processing method provided in some embodiments of this application; Figure 2 This is a second flowchart of a data processing method provided in some embodiments of this application; Figure 3 This is the third flowchart of a data processing method provided in some embodiments of this application; Figure 4 This is a flowchart of a data processing method provided in some embodiments of this application; Figure 5 This is the fifth flowchart of a data processing method provided in some embodiments of this application; Figure 6 This is a structural block diagram of a data processing apparatus provided in some embodiments of this application; Figure 7This is a schematic diagram of the structure of an electronic device provided in some embodiments of this application; Figure 8 This is a schematic diagram of the hardware structure of an electronic device that implements the various embodiments of this application. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0013] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0014] The data processing method provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0015] Figure 1 This is one of the flowcharts of a data processing method provided in some embodiments of this application, applied to a first terminal, such as... Figure 1 As shown, the method may include the following steps: step 101, step 102 and step 103.
[0016] In step 101, the first terminal generates a local knowledge index based on local data.
[0017] In this embodiment, both the first terminal and the second terminal cover a variety of common terminal devices such as smartphones, tablets, smart wearable devices, and desktop computers. This method has wide applicability and can be implemented on a variety of everyday electronic devices.
[0018] In this embodiment of the application, the first terminal performs in-depth content analysis on the locally stored data (i.e., local data) to mine key information and relationships in the local data, and generates a local knowledge index based on this.
[0019] In this embodiment of the application, the local knowledge index is usually a directory of local data that is rich in content and not anonymized. It organizes the important features and key points of the local data in a structured way, which is equivalent to creating a detailed internal archive for all the data on the first terminal, so as to facilitate more efficient management and searching of the data in the future.
[0020] As can be seen, in this embodiment, the local knowledge index is a digital archive of user data stored locally on the terminal. It contains a multi-dimensional understanding of the original data, while the original data itself remains in its original state. The content of the local knowledge index is in a non-anonymized state, providing a complete data foundation for local AI analysis and subsequent anonymization operations.
[0021] In some embodiments, step 101 may specifically include the following step: step 1011.
[0022] In step 1011, the local data is processed by a natural language processing model to perform entity recognition, keyword extraction, and semantic vectorization to generate a local knowledge index containing semantic information.
[0023] In this embodiment of the application, the natural language processing model can scan and analyze local data to identify various entities. These entities can be specific names of people, places, organizations, or special codes, or abstract concepts and events. By identifying entities, the key elements in the local data can be sorted out, providing a data foundation for subsequent understanding and analysis.
[0024] In this embodiment, the natural language processing model can extract keywords from local data that summarize the data's theme and core content. These keywords are key features of the data and can concisely express its main information. Keyword extraction helps to quickly understand the general content of the data, facilitating subsequent searching and classification.
[0025] In this embodiment, the natural language processing model can map the identified entities, extracted keywords, and the semantic information of the entire text into a high-dimensional numerical space, generating corresponding semantic vectors. This is the process of converting text data into numerical vectors. Semantic vectors can capture the semantic similarity and correlation between texts. Even if two texts are expressed differently, as long as their semantics are similar, their semantic vectors will be relatively close, thereby realizing the transformation of text meaning into a computable distance relationship.
[0026] In this embodiment, semantic vectorization and other technologies enable the terminal to understand the inherent meaning and concepts of the data, rather than relying solely on literal keyword matching, thereby improving the accuracy and intelligence of subsequent searches.
[0027] It can be seen that in the embodiments of the present application, by integrating the results of entity recognition, keyword extraction, and semantic vectorization processing, a local knowledge index containing semantic information is generated. This index not only contains explicit information such as entities and keywords in the data, but also contains the semantic features of the data, and can more comprehensively and accurately reflect the content and structure of the local data, providing a data basis for subsequent construction of a global data graph and efficient and accurate cross-device semantic search.
[0028] In some embodiments, agents can be deployed on multiple terminals of the user (such as mobile phones and computers). After obtaining authorization, the agents on each terminal execute the steps of step 101 to step 103. Taking step 101 as an example, the agent in the first terminal performs in-depth content analysis and semantic understanding on files, application data, etc. locally, and generates a local knowledge index.
[0029] In the embodiments of the present application, for the agent deployed in the first terminal, it can perceive data changes in real time by listening to file system events of the operating system (such as creation, modification, deletion operations) or by calling the APIs of specific application programs (such as note application programs, email client application programs). Different from traditional software that only records file names and sizes, it uses built-in lightweight AI models (such as NLP, CV models, etc.) to deeply mine the file content. It will perform operations such as entity recognition (finding names of people, places, and organizations in the text), keyword extraction (grabbing core topic words), semantic vectorization (converting the meaning of text into a digital vector), and structured information extraction (such as identifying amounts and dates from bill pictures), and finally generate a rich and non-desensitized local knowledge index.
[0030] Exemplarily, taking a Word document named "Phoenix Project Contract_v3.docx" as an example, the entry corresponding to this document in the local knowledge index contains the following content: 1) Core location information: data_pointer: C: / Users / ZhangSan / Projects / Phoenix / Phoenix Project Contract_v3.docx, used to locate the storage path of the document in the device; handler_application: Microsoft Word, used to identify the application program for processing this document; 2) Basic structure information: file_name: Phoenix Project Contract_v3.docx, which is the name of the document; file_type: DOCX, which is the type of the document; modification_date (modification date): 1753090155, records the date the document was modified; 3) Content-derived semantic information (local, non-anonymized version): semantic_vector_embedding: [-0.56, 0.81, ..., 0.33], is a vector representing the core meaning of the document; extracted_keywords: [“Contract”, “Software Development”, “Confidentiality Agreement”, “Penalty”, “Phoenix Project”], which are keywords extracted from the document; named_entities:{“PERSON”:[“Zhang San”,“Li Si”],“ORG”:[“Company A”,“Company B”],“MONEY”:[“500000 yuan”,“600000 yuan”,“700000 yuan”],“PERCENT”:[“5%”,“10%”],“PROJECT_CODE”:[“Phoenix Project”]}, represents named entities identified in the document, including people, organizations, amounts, percentages, project codes, etc. language:zh-CN, identifies the language used in the document.
[0031] In step 102, the first terminal performs de-identification processing on the local knowledge index to obtain de-identified metadata; wherein, the metadata includes descriptive information and location information.
[0032] In this embodiment of the application, since the local knowledge index contains original information, in order to ensure information security, the local knowledge index cannot be directly uploaded to the server. Therefore, the first terminal will perform desensitization processing on the local knowledge index to remove or hide the parts that may involve user privacy or sensitive information.
[0033] In this embodiment of the application, the descriptive information is the de-identified semantic information used to provide a general description of the data, such as the hashed entity, semantic vector, file type, topic, structural summary, etc.
[0034] In this embodiment, the location information is a security identifier pointing to the local original data, which is used to accurately locate the data when subsequent operations are needed, while avoiding the exposure of the real path structure.
[0035] In this embodiment, a local knowledge index containing rich details is transformed into metadata that has undergone strict anonymization processing and can be securely reported to the server. This ensures that the information stored on the server can support efficient global search while preventing the decryption of any user's private content, thus protecting user privacy and data security while ensuring data availability.
[0036] In some embodiments, different desensitization methods can be used for different types of sensitive information. Accordingly, step 102 may specifically include at least one of the following steps: step 1021, step 1022 and step 1023.
[0037] In step 1021, an irreversible hash operation is performed on the identified named entities to generate hash values.
[0038] In this embodiment of the application, the identification of named entities in the text has usually been completed in the early stage of the data processing flow. These named entities cover various entities with clear meanings, such as personal names, place names, organization names, project codes, and specific terms.
[0039] In this embodiment, hashing is a process of converting input information of arbitrary length (referring to named entities in this case) into fixed-length output information (i.e., hash values) using a specific hash algorithm. Irreversibility means that the original named entity cannot be derived from the generated hash value. For example, using the common MD5 hash algorithm, hashing "Zhang San" will produce a fixed-length hash value "81dc9bdb52d04dc20036dbd8313ed055", but "Zhang San" cannot be retrieved from this hash value.
[0040] In this embodiment, generating hash values can be used to protect the privacy of named entities while retaining some identifying characteristics for data analysis and association. For example, in data sharing scenarios, different organizations can determine whether the same entity is involved by comparing hash values without knowing the entity's specific name.
[0041] In step 1022, sensitive information with a preset format is identified and filtered out.
[0042] In this embodiment, the preset format refers to a pattern pre-defined based on the characteristics of different types of sensitive information. For example, ID card numbers typically have a fixed 18-digit combination format, bank card numbers also have specific digit counts and arrangement rules, telephone numbers have corresponding format specifications in different regions, and email addresses also have specific naming rules.
[0043] In this embodiment, pattern matching techniques such as regular expressions can be used to scan and match the content in text data to determine whether there is sensitive information that conforms to a preset format. For example, the regular expression \d{17}[\dXx] can be used to identify information that conforms to the format of an ID card number.
[0044] In this embodiment, if sensitive information conforming to a preset format is identified, it is removed from the data or replaced with a specific placeholder to protect the sensitive information. For example, in a report containing a user's ID card number, the ID card number is replaced with "[ID card number has been de-identified]". Another example is processing a phone number as a hash ("138XXXX5678") to support specific searches while maximizing information security.
[0045] In step 1023, the specific numerical information is replaced with an abstract type identifier or range quantization value.
[0046] In this embodiment of the application, the specific numerical information may include various types of numeric data, such as age, income, amount, percentage, number of shares, quantity, etc. For example, a personal financial report may include specific values such as an individual's monthly income of "8,500 yuan" and savings amount of "50,000 yuan".
[0047] In this embodiment, replacing specific numerical values with abstract type identifiers means classifying the values into a predefined category based on their characteristics. For example, age "25 years old" is replaced with "youth," and income "8500 yuan" is replaced with "middle income." Alternatively, specific numerical values can be replaced with range-quantified values, that is, using a numerical range to represent the original value. For example, deposit amount "50000 yuan" is replaced with "40000-60000 yuan," and product sales "1250 units" is replaced with "1000-1500 units."
[0048] In this embodiment of the application, a secondary filtering can be performed to remove all identified named entities, specific values, and other sensitive information, retaining only the macro-structural description of the document. For example, for a contract, the desensitized summary might be "This is a legal document containing 5 main chapters and approximately 3,000 words, defining the rights and obligations of Party A and Party B," without mentioning who Party A and Party B are, or the specific content of their rights or obligations.
[0049] In this embodiment, a configurable, user-customizable stop word list or sensitive word filter can be used. This stop word list or sensitive word filter will by default include common highly sensitive words such as passwords and internal project codenames, ensuring that they are not extracted and uploaded as ordinary keywords.
[0050] As can be seen, in this embodiment, the original named entity information is hidden through irreversible hashing. Even if the data is leaked during transmission or sharing, attackers cannot obtain the real named entity, thus effectively protecting the privacy of individuals or organizations. Timely identification and filtering of sensitive information with preset formats can reduce the presence of sensitive information in data at the source, lowering the security risks caused by data leakage. Replacing specific numerical information with abstract type identifiers or range quantification values protects data privacy while preserving a certain degree of data usability.
[0051] In some embodiments, the location information of the metadata may include: a data pointer.
[0052] In this embodiment, a data pointer is a stable and secure "address" or "reference" that points to local data; it is essentially an identifier. It differs from the original, human-readable file storage path, such as "C: / Users / My Documents / contract.docx." Using this non-original path format ensures data privacy and security. In a computer storage architecture, data can be distributed across different storage devices, folders, or database tables. A data pointer acts like a precise coordinate, enabling quick location of the required data.
[0053] In this embodiment of the application, the data pointer specifically includes the following key elements: The device_id serves as a unique identifier for the source device, used to clearly identify the origin of the data. For example, it can be set to "Zhang San's MacBook Pro" to ensure accurate differentiation of data sources in a multi-device environment. The `data_pointer` serves as a stable pointer to local data and has two implementations: First, it uses the hash value of the file path. This involves applying a specific hash algorithm to a file path like "hash(C: / Users / ...)" to generate a string of characters without actual semantic meaning, such as "a1b2c3d4...". When receiving instructions, the local terminal can locate the specific file based on this hash value, effectively avoiding exposing the original file path structure on the server side and ensuring data privacy and security. Second, for non-file data, it can use an internal application data ID. This ID is unique within the application and can accurately point to specific non-file data.
[0054] The handler_application is used to identify the application required to process the data. For example, it can be set to "Microsoft Word", "Notion", etc., so that the correct application can be called to perform operations during data processing.
[0055] As can be seen, in this embodiment, the data pointer provides an efficient data location mechanism, enabling the device to quickly find the required data, reducing data search time, and improving response speed and performance. This ability to quickly locate data is particularly important when processing large amounts of data. At the same time, the data pointer also effectively protects the privacy information of users' local file storage.
[0056] In some embodiments, the descriptive information of the metadata may include at least one of the following: infrastructure information and content-derived semantic information.
[0057] In this embodiment, the infrastructure information is used to describe the basic physical and structural properties of the data. It provides detailed information about the physical storage and structural organization of the data, helping users and devices understand the basic characteristics and storage methods of the data.
[0058] In this embodiment, the infrastructure information may include: file type, file size, creation time, and modification time. The file type specifies the file format in which the data exists, such as text files (.txt), image files (.jpg, .png), and audio files (.mp3). Different file types have different encoding methods and processing requirements; understanding the file type helps in selecting the appropriate application to open and process the data. The file size indicates the amount of storage space occupied by the data, typically measured in bytes (Byte), kilobytes (KB), megabytes (MB), gigabytes (GB), etc. File size information is crucial for assessing storage needs, data transfer time, and resource consumption. The creation time and modification time record the timestamps of the data's creation and last modification; this time information can be used to track the data's change history and manage data versions.
[0059] For example, infrastructure information may specifically cover the following: file_name is the original filename, for example, "Phoenix Project Contract_v3.docx". This information is considered low-sensitivity information and is mainly used for display in the user interface to facilitate user identification and retrieval of files.
[0060] file_type / mime_type are used to specify the file type, such as "DOCX", "PDF", "image / jpeg", etc., so that the device can select the appropriate processing method according to the file type.
[0061] file_size, creation_date, and modification_date represent the file size, creation timestamp, and modification timestamp, respectively. This information can be used to sort and filter files, improving the efficiency of file management.
[0062] In this embodiment, content-derived semantic metadata is generated by the terminal after analyzing the original content, and it possesses highly abstract and non-sensitive semantic information. Content-derived semantic information is used to describe the content of the data; it not only focuses on the surface form of the data but also delves deeper into the semantics and meaning contained within the data, helping users and devices better understand the data's theme, key information, and inherent relationships.
[0063] In this embodiment of the application, content-derived semantic information may include at least one of the following: semantic vector embedding, hashed named entity list, and structural summary.
[0064] Semantic vector embedding, for example, converts data content into vector form, which captures the semantic features of the data. By calculating the similarity between vectors, operations such as semantic search, classification, and clustering of the data can be achieved. For instance, in natural language processing, word embedding techniques can be used to convert words in text into vectors, and then the semantic similarity between words can be determined by calculating the distance between the vectors. Taking a hashed list of named entities as an example, named entities are entities in text that have specific meaning, such as person names, place names, and organization names. Hashing named entities to generate hash values can protect their privacy while retaining their identifying characteristics for data analysis and association. The hashed list of named entities can be used to quickly identify and extract key entities from text, enabling entity relationship analysis and graph construction. Taking structured summaries as an example, data content is extracted and summarized in a structured manner to generate concise summary information. This summary information highlights the main points, key data, and important conclusions of the data, helping users quickly understand the core content and improving information retrieval efficiency. For example, for a long report, a structured summary containing main chapter titles and key conclusions can be generated.
[0065] For example, content-derived semantic metadata may specifically cover the following: Semantic vector embedding refers to the process of calculating a numerical vector, such as a 1536-dimensional floating-point array, from the semantic content of a document as a whole or key paragraphs using a model. This vector represents the meaning of the content and can be used for "concept search," such as searching for "disclosure agreement" to find files named "NDA." Furthermore, the vector itself does not contain any readable text information, effectively protecting data privacy.
[0066] `extracted_keywords` refers to common keywords and topic terms extracted from text by the local agent, such as "contract," "software development," and "financial report." During the extraction process, the local agent uses a configurable stop word list or sensitive word filter to ensure that highly sensitive words such as project codes and passwords are not uploaded as keywords, further ensuring data security.
[0067] `named_entities_hashed` refers to the hash values of named entities (such as person names, company names, and place names) identified by the local agent. Instead of uploading the original text, the hash values of these entities are uploaded. For example, hashing "Zhang San" using the SHA256 algorithm yields "a665a4...". The server only knows that all documents containing "a665a4..." are related to the same entity, but cannot decipher the name "Zhang San". This method achieves associative search while protecting the privacy of entity information.
[0068] Abstract summary refers to an automatically generated structured summary that does not contain specific details. For example, for a contract, the summary might be "This is a legal document containing 5 main chapters and about 3,000 words, defining the rights and obligations of Party A and Party B," without mentioning the specific rights or obligations. This provides basic information about the document while protecting its sensitive content. Language refers to the primary language used to identify content, such as "zh-CN" or "en-US," so that intelligent agents can process and display content accordingly based on language information.
[0069] For the first terminal that has deployed the intelligent agent, the intelligent agent performs desensitization processing on the local index information and extracts only the metadata.
[0070] In step 103, the first terminal uploads metadata to the server; wherein the server constructs a global data graph based on metadata from at least two terminals.
[0071] In this embodiment of the application, after the processing of steps 101 to 103, all sensitive information processing is completed at the data source (i.e., the user terminal local), ensuring privacy and security.
[0072] In this embodiment, the de-identified metadata is uploaded to a server, enabling this data to enter a broader storage and management environment. The server has powerful aggregation capabilities, receiving metadata from multiple different terminals and integrating these scattered data. Through integration, a global data graph is constructed, depicting the distribution and relationships of user data across different terminals at a macroscopic level. The global data graph plays a crucial role, achieving logically unified management of user-wide distributed data. Regardless of which terminals the data is stored on, it can be uniformly managed and scheduled through this graph. Simultaneously, it ensures secure data interaction, guaranteeing that sensitive user information is not leaked during data sharing and use, providing users with a secure and efficient data management environment.
[0073] For example, for a contract document, only its file name, modification date, device location, and named entities, keywords, and semantically representative vector embeddings extracted from the content and hashed are reported, while the original content and specific values of the contract are not uploaded.
[0074] In this embodiment of the application, the server is responsible for collecting metadata from all terminals and constructing a global data graph that describes the distribution and association of user's global data. This graph itself does not contain any original file content, but only serves as a directory of the entire knowledge base, providing users with a global data perspective and facilitating data management and search operations.
[0075] In some embodiments, taking a first terminal with an intelligent agent deployed as an example, the intelligent agent first performs deep content analysis and semantic understanding on the user-authorized files and application data locally, generating a detailed local knowledge index. This process is entirely performed locally on the first terminal, and the raw data does not leave the first terminal. Then, the intelligent agent performs refined anonymization and abstraction processing on the local knowledge index, only synchronizing the processed structured metadata to the server through an end-to-end encrypted channel. Finally, the server aggregates the metadata from all terminals to construct a global data graph describing the distribution and relationships of user-wide data.
[0076] As can be seen from the above embodiments, in this embodiment, on the one hand, the terminal performs de-identification processing on the local knowledge index, effectively removing sensitive information that may leak user privacy, and performs security protection before the data is uploaded to the server. In this way, even if the metadata is illegally obtained during transmission or storage, attackers cannot obtain the user's sensitive information from it, reducing the risk of data leakage. On the other hand, users do not need to manage data on different terminals separately. They can search the entire data set simply by using the global data graph on the server side, which improves the efficiency of information retrieval in multi-terminal device scenarios.
[0077] Figure 2 This is a second flowchart of a data processing method provided in some embodiments of this application, applied to a first terminal, such as... Figure 2 As shown, the method may include the following steps: step 201, step 202, step 203, step 204, step 205, step 206 and step 207.
[0078] In step 201, the first terminal generates a local knowledge index based on local data.
[0079] In step 202, the first terminal performs de-identification processing on the local knowledge index to obtain de-identified metadata; wherein, the metadata includes descriptive information and location information.
[0080] In step 203, the first terminal uploads metadata to the server; wherein the server constructs a global data graph based on metadata from at least two terminals.
[0081] The content of steps 201 to 203 in the embodiments of this application is the same as that of... Figure 1 The steps 101 to 103 in the illustrated embodiment are similar and will not be repeated here.
[0082] In step 204, the first terminal receives the user's search request.
[0083] In this embodiment of the application, the user initiates a search request through the interactive interface of a first terminal (such as a mobile phone, computer, or other device). The search request can express the user's intent in the form of natural language or structured search conditions, including but not limited to the following two situations: the user intends to directly obtain specific information content stored in other terminals, and the user intends to locate a terminal device containing the target information.
[0084] For example, when a user enters "find the document I wrote last week about the Phoenix project," their core intent might be to directly obtain the document content or to determine the location of the terminal storing the document. After receiving the search request, the first terminal will process it uniformly, complete semantic matching and terminal location in the global data graph through subsequent steps, and return the corresponding location results to the user.
[0085] In step 205, the first terminal parses and de-identifies the user's search request to generate an encrypted search request.
[0086] In this embodiment of the application, in order to protect user privacy, the first terminal will not send the user's original search request directly to the server, but will first perform security processing on the search request locally, which may include the following steps: Parsing and processing: The local natural language processing engine understands the user's search intent and identifies key entities (such as Phoenix Project, last week), operation type (such as search) and constraints; Desensitization: Sensitive entities identified are processed using the same desensitization algorithm as those in the data indexing stage (e.g., hashing "Phoenix Project") to ensure that the search criteria received by the server and the stored metadata are under the same privacy protection. Semantic vectorization: converting the semantic content of a search request into a search vector; Request encapsulation: The processed search conditions, search vectors, and other search parameters are encapsulated into a structured data packet and encrypted to obtain an encrypted search request.
[0087] In this embodiment, the search vector is a high-dimensional numerical vector (e.g., an array of hundreds of floating-point numbers). It can be generated by a local AI model after performing deep semantic analysis on the semantic content of the user's entire search request (e.g., "find the document I wrote last week about project planning"). This vector is used to convert the user's vague and abstract natural language request into a precise mathematical representation that can be understood and computed by a computer. It captures the semantic essence of the search request rather than the surface words.
[0088] In this embodiment, the server can find files that are closest in meaning to the user's search by calculating the similarity (such as cosine similarity) between the search vector and the semantic vector of the metadata of each file in the global data graph. Even if the title or content of these files does not contain the exact keywords searched by the user, they can still be found.
[0089] In this embodiment, the constraints are a set of standard filtering rules based on metadata, which may include at least one of time range, file type, keywords and hashed entity identifiers, and are used to perform preliminary screening of metadata before semantic matching.
[0090] In this embodiment of the application, the global data graph may contain tens of thousands or even hundreds of thousands of data nodes. If the similarity between the search vector and all nodes is directly calculated, the computational cost will be very high. However, the constraints (such as time_range and file_type) can first filter out "PDF files from last week" at a very fast speed (e.g. based on inverted index), reducing the candidate set from tens of thousands to dozens. Then, fine semantic matching is performed on this small set, which reduces the computational burden and ensures the response speed.
[0091] It's important to note that search vectors and constraints do not work in isolation; rather, they collaborate through a hybrid search process that prioritizes filtering over semantics. Specifically, in the first stage, the server performs a fast and precise initial screening using constraints—a simple "yes or no" Boolean filtering process that significantly narrows the search scope and generates a manageable set of candidate metadata nodes. In the second stage, the server then calculates the similarity score between the search vector and the semantic vector of each candidate node, ranking the candidates based on relevance and placing the results that best match the user's semantic intent first. In the third stage, the final results are sorted according to semantic similarity scores, and may be fine-tuned based on other factors (such as timeliness) before being returned.
[0092] As can be seen, in this embodiment of the application, by combining multiple constraints, the server can first perform a fast preliminary screening, which greatly narrows the search range (such as quickly screening dozens of candidate nodes from tens of thousands of nodes), reduces the overhead for subsequent computationally intensive semantic similarity matching, and improves the overall response speed and scalability.
[0093] In step 206, the first terminal sends an encrypted search request to the server; wherein, the server obtains the user's search request based on the encrypted search request, determines the second terminal where the search data corresponding to the user's search request is located based on the global data map, and returns the location result to the first terminal; the location result includes at least the device identifier of the second terminal.
[0094] In this embodiment of the application, the first terminal sends the encrypted search request to the server through a secure channel (such as TLS, SSL, etc.), wherein the search request is used to instruct the server to perform a search in the global data graph.
[0095] In some embodiments, in addition to including the device identifier of the second terminal, the location result may also include at least one of the following: file name, file type, and modification date.
[0096] In this embodiment, the core function of the location result is to guide the user to the location of the data. If only the device ID (such as "PC_B") is returned, the user will be forced to perform a secondary operation: they still need to remotely connect or log in to a second terminal and search through a potentially massive number of files to locate the target file. However, by returning key attributes such as file name and file type, the user can directly confirm the file they are looking for on the first terminal (such as confirming it by the file name Phoenix Project Final Contract.docx), without performing any additional connection or search operations, thus achieving a seamless experience from search to confirmation.
[0097] In step 207, the first terminal receives the location result returned by the server.
[0098] In this embodiment, after the server completes the map search, it does not return any original data content, but only a location result. This location result is essentially a navigation instruction, which at least includes a target device identifier (second terminal ID) to inform the first terminal which device the data it is looking for is located on (such as "your office computer"). In addition, it usually includes basic metadata such as file name and type for result display.
[0099] For example, when a user wants to find contracts related to "Project Phoenix", they can enter a search request on the intelligent agent interface of mobile phone A, such as "find the contract drafts related to 'Project Phoenix' from last week". The intelligent agent will not directly send the raw text to the server, but will perform preprocessing operations locally to generate a structured encrypted search request. The specific steps are as follows: Intent and Entity Recognition: The locally deployed AI analysis engine performs natural language understanding processing on the user's input. Through natural language understanding algorithms, it determines that the core intent of the user's input is "search for files". Using natural language parsing technology, key entities and constraints are extracted from the input, specifically including: the time range "last week", the project code / keyword "Phoenix Project", and the file type clue "contract draft". Entity desensitization and vectorization: For the identified entity "Phoenix Project", the same hashing algorithm as in the data indexing stage is used to perform a hash operation to generate a hash value hash("Phoenix Project")->c3b8... This hash value is used for subsequent precise matching operations to ensure the security and privacy of entity information during transmission and processing.
[0100] Search statement vectorization: Using the local AI engine, the overall semantics of the search statement "find the contract draft about 'Phoenix Project' from last week" are calculated and transformed into a search vector. This search vector can accurately represent the core meaning of the user's search and provide a foundation for subsequent semantic matching in the global data graph.
[0101] Constructing the search data packet: The agent encapsulates the results obtained after intent and entity recognition, entity desensitization, and vectorization into an encrypted, structured search request. Taking JSON format as an example, the search request is as follows: JSON { "user_id": "user_alpha", "query_vector": [-0.56, 0.81, ..., 0.33], "filters": { "date_range": { "start": 1752403200, "end": 1753007999}, "keywords_contain": ["contract", "draft"], "hashed_entities_contain": ["c3b8..."] }, "top_n_results": 5 }
[0102] As can be seen from the above embodiments, this embodiment achieves cross-device data discovery with privacy and security by using localized parsing and de-identification technologies, ensuring that the user's original search intent and sensitive information do not leave the terminal. The terminal uses semantic vector matching technology to accurately understand the user's natural language intent and completes accurate searches through lightweight encrypted data interaction, reducing the risk of privacy leakage while ensuring data transmission efficiency. The finally returned device identification information provides a precise routing foundation for subsequent secure remote operations, constructing a distributed data management paradigm that integrates privacy protection, semantic understanding, and efficient interaction.
[0103] Figure 3 This is a flowchart of a data processing method provided in some embodiments of this application, applied to a first terminal, such as... Figure 3 As shown, the method may include the following steps: step 301, step 302, step 303, step 304, step 305, step 306, step 307, step 308, step 309, step 310 and step 311.
[0104] In step 301, the first terminal generates a local knowledge index based on local data.
[0105] In step 302, the first terminal performs de-identification processing on the local knowledge index to obtain de-identified metadata; wherein, the metadata includes descriptive information and location information.
[0106] In step 303, the first terminal uploads metadata to the server; wherein the server constructs a global data graph based on metadata from at least two terminals.
[0107] The content of steps 301 to 303 in the embodiments of this application is the same as that of... Figure 1 The steps 101 to 103 in the illustrated embodiment are similar and will not be repeated here.
[0108] In step 304, the first terminal receives the user's search request.
[0109] In step 305, the first terminal parses and de-identifies the user's search request to generate an encrypted search request.
[0110] In step 306, the first terminal sends an encrypted search request to the server; wherein, the server obtains the user's search request based on the encrypted search request, determines the second terminal where the search data corresponding to the user's search request is located based on the global data map, and returns the location result to the first terminal; the location result includes at least the device identifier of the second terminal.
[0111] In step 307, the first terminal receives the location result returned by the server.
[0112] The content of steps 304 to 307 in the embodiments of this application is the same as that of... Figure 2 Steps 204 to 207 in the illustrated embodiment are similar and will not be repeated here.
[0113] In step 308, the first terminal receives an operation instruction for remote target data on the second terminal; wherein the positioning result is used to point to the remote target data.
[0114] In this embodiment of the application, after obtaining the location of the target data from the positioning results, the user issues a specific operation instruction on the first terminal (such as a mobile phone). The operation instruction is usually in natural language (such as "change the liquidated damages in this contract from 5% to 10%)" or GUI operation.
[0115] In step 309, the first terminal obtains the operation intent data packet according to the operation instructions.
[0116] In this embodiment of the application, in order to achieve secure remote operation, the first terminal does not record the screen or transmit files. Instead, it uses natural language processing technology to decompose the operation instructions into operation intent (such as UPDATE_CONTENT) and a series of parameters (such as search_context: "penalty", old_value: "5%", new_value: "10%)). The intent and parameters are then encapsulated into a structured (such as JSON format) machine-readable operation intent data packet.
[0117] In some embodiments, the operation intent data packet is a JSON object that defines the operation type and operation parameters; wherein, the operation type may include at least one of the following: content update, file deletion, file read, file rename; and the operation parameters may include at least one of the following: location information of remote target data, search context, original value to be verified, and new value to be set.
[0118] In some embodiments, in order to ensure the security of data transmission, the following steps may be included before step 310 above: encrypting the operation intent data packet using the public key or pre-shared symmetric key of the second terminal, and digitally signing the encrypted data packet using the private key of the first terminal.
[0119] In step 310, the first terminal sends the operation intent data packet to the second terminal; wherein, the second terminal parses the operation intent data packet, performs the operation defined in the operation intent data packet on the remote target data, and returns the operation execution result to the first terminal.
[0120] In this embodiment, the first terminal sends a lightweight operation intent data packet directly to the second terminal where the target data is located via a secure channel. What is transmitted here is not the data itself, but rather a command to perform an operation on the data.
[0121] In step 311, the first terminal receives the operation execution result returned by the second terminal.
[0122] In this embodiment, after the second terminal completes the operation on the target data locally, it returns an operation execution result to the first terminal. This operation execution result is an encrypted and signed status confirmation message. The status confirmation message includes a status code indicating success or failure of the operation, as well as necessary error information (such as {“status”: “SUCCESS”}), but does not contain the original content of the remote target data. In practical applications, the operation execution result is typically a lightweight status confirmation message.
[0123] For example, the original contract is stored on PC B. If the user's intention is to modify the contract, such as issuing a command on mobile phone A: "Change the penalty for breach of contract from 5% to 10%", the agent on mobile phone A will not request the transmission of the contract. Instead, it will parse the command and convert it into a structured, lightweight action intent data packet (such as a JSON object). This encrypted action intent data packet is then sent to PC B over the network.
[0124] Specifically, the intelligent agent on phone A uses its local natural language understanding engine to perform deep parsing of the operation commands. First, it identifies the user's core intent as UPDATE_CONTENT (content update) and extracts all the parameters (i.e., "slots") required to execute this intent. In this example, the extracted slots include: target_file_context: "Phoenix Project Contract", used to reconfirm the operation objective; `search_context: "penalty"` is used to locate the modified area within the document; old_value: "5%" is used for pre-execution verification to ensure the accuracy of the operation; new_value: "10%" is used to indicate new content that is about to be updated.
[0125] Next, the agent combines the parsed intent and slot with the location results to construct a structured operation intent data packet. This data packet is the core carrier for secure remote operation and is typically organized in JSON format, containing a unique ID for this operation, the target device and file pointers, the specific actions, and parameters. An example is shown below: JSON { "intent_id": "op-uuid-67890", / / Unique identifier for this operation "target_device": "PC_B", / / Target device ID "target_pointer": "hash_of_path_to_contract.docx", / / Pointer to the target file "action": "UPDATE_CONTENT", / / Operation type "parameters": { "search_context": "penalty", "old_value_to_verify": "5%", "new_value_to_set": "10% }, "timestamp": 1753090100 } Before transmission, to ensure security, the JSON data packet undergoes two processing steps: The entire JSON object is encrypted using the target device's (PC_B) public key or a pre-shared symmetric key (such as AES-256) to ensure confidentiality. The encrypted data packet is then signed using the requesting device's (phone A) private key to ensure the authenticity and non-repudiation of the instruction, preventing forgery or tampering.
[0126] Finally, this encrypted and signed, lightweight data packet is sent from phone A to PC B through an established secure channel that also uses TLS encryption.
[0127] As can be seen from the above embodiments, this embodiment can reduce the risk of remote operation data leakage, only transmit the operation intent, and keep the original data locally on the second terminal to avoid being stolen or spied on during transmission; in addition, the data volume of the operation intent data packet is small, which can achieve efficient low bandwidth and high response interaction, and can respond instantly even in weak networks, far exceeding the traditional solution of transmitting pixel streams or complete files.
[0128] Figure 4 This is a flowchart of a data processing method provided in some embodiments of this application, applied to a server, such as... Figure 4 As shown, the method may include the following steps: step 401 and step 402.
[0129] In step 401, the server receives anonymized metadata uploaded by at least two terminals; wherein the metadata includes: description information and location information.
[0130] In this embodiment of the application, the server receives encrypted metadata from multiple terminals authorized by the user (such as mobile phones, tablets, and PCs). This metadata has undergone strict desensitization processing on each terminal and does not contain any original data content or sensitive information that can be decrypted.
[0131] In step 402, the server constructs a global data graph based on metadata from at least two terminals.
[0132] In this embodiment, the server does not simply store this metadata, but rather associates them to construct an interconnected, semantically rich global data graph. In this graph, each metadata element is a node, and nodes are connected through shared hash entities (such as the same hash value representing the same person), thereby forming an inherent relationship between data scattered across different terminals.
[0133] In this embodiment of the application, the global data map acts as a "unified index" or "global map" for all user data, allowing users to logically and transparently access, search and manage data on all terminals as if they were in one place, while in reality the original data is still physically distributed and stored on their respective local terminals.
[0134] In this embodiment, the global data graph itself is composed of anonymized metadata, ensuring that all search operations are performed only at the metadata level. The server does not need to access, and cannot access, any user's original data content, thus guaranteeing the privacy and security of the search process from an architectural perspective.
[0135] In this embodiment, the semantic vectors and hash entities stored in the global data graph are the foundation for supporting cross-device semantic search and related search. It can understand the user's search intent and discover deep relationships between data (such as "find all documents related to 'Zhang San'", even if these documents are scattered across different terminals). In this embodiment, the server processes lightweight metadata in a way that saves far more computing and storage resources than synchronizing and indexing massive raw data files. This design can be efficiently scaled to support a large number of users and terminals while maintaining fast response capabilities.
[0136] As can be seen from the above embodiments, in this embodiment, on the one hand, users do not need to manage data on different terminals separately; they can search the entire data set simply by using the global data graph on the server side, improving the efficiency of information retrieval in multi-terminal device scenarios. On the other hand, the local knowledge index is anonymized, effectively removing sensitive information that may leak user privacy. Security protection is implemented before the data is uploaded to the server, so even if metadata is illegally obtained during transmission or storage, attackers cannot obtain sensitive user information from it, reducing the risk of data leakage. This effectively breaks down barriers between devices, enabling unified management, efficient search, and secure utilization of distributed data.
[0137] Some embodiments provided in this application, the data processing methods provided, in Figure 4 After step 402 of the illustrated embodiment, the following steps may also be included: step 403, step 404 and step 405.
[0138] In step 403, the server receives an encrypted search request sent by the first terminal.
[0139] In this embodiment of the application, the encrypted search request packet is a pre-processed structured request that contains relevant information such as search vectors and constraints (such as time range, file type, hashed entity, etc.) that can accurately describe the user's search needs. It exists in encrypted form to prevent it from being illegally stolen or tampered with during transmission.
[0140] In step 404, the server obtains the user's search request based on the encrypted search request, determines the second terminal where the search data corresponding to the user's search request is located based on the global data map, and obtains the location result; wherein, the location result includes at least the device identifier of the second terminal.
[0141] In this embodiment, after receiving an encrypted search request from the first terminal, the server utilizes its powerful computing capabilities and stored global data graph to perform a search operation. The global data graph is a comprehensive data index structure that integrates data information from various terminals, including data type, content characteristics, and storage location. The server decrypts and analyzes the encrypted search request to obtain the user's search request, then extracts key search conditions, performs precise matching and searching within the global data graph, thereby locating the second terminal containing the data matching the search request and obtaining relevant location result information such as the device identifier of that second terminal.
[0142] In step 405, the server returns the location result to the first terminal.
[0143] In this embodiment, after obtaining the location result, the server returns this crucial information to the first terminal. This location result typically includes the device identifier of the second terminal, as well as low-sensitivity metadata such as filename and file type, but not the original data content itself. In this way, the first terminal can determine which terminal the target data is stored on based on the returned location result, thus providing a basis for subsequent possible searches, operations, and other actions.
[0144] As can be seen from the above embodiments, in this embodiment, the server processes encrypted requests and anonymized metadata throughout the entire process. It cannot know the user's original search intent, nor can it see any of the user's actual data content, thus avoiding the risk of privacy leakage and achieving true blind search. Furthermore, the returned location results are lightweight location information (device ID, filename, etc.) rather than a large raw data file, saving network bandwidth. The location results clearly indicate the physical location of the data (second terminal), securely transferring the responsibility and permissions for data operations back to the user's device, providing an accurate access address for subsequent secure remote operations, while the server itself does not bear the security risks associated with directly manipulating the data.
[0145] In some embodiments provided in this application, step 404 may specifically include the following steps: step 4041, step 4042, step 4043, step 4044 and step 4045.
[0146] In step 4041, the server decrypts the encrypted search request and extracts the search vector and constraints.
[0147] In this embodiment of the application, after receiving the encrypted search request sent by the first terminal, the server needs to decrypt it first. The decryption operation can use a specific encryption algorithm and key to restore the encrypted data into readable information.
[0148] In this embodiment, the extracted search vector is a digital representation of the user's search intent. It uses a specific algorithm to convert the user's input search content (such as text, keywords, etc.) into vector form for subsequent semantic matching and calculation.
[0149] In this embodiment, constraints are various restrictions set by the user during the search, such as data type (documents, images, videos, etc.), time range (searching for data within a specific time period), data size, etc. These constraints are used to further narrow the search scope and improve the accuracy of the search.
[0150] In step 4042, based on constraints, the metadata nodes in the global data graph are initially screened to obtain a set of candidate metadata nodes; wherein each metadata node contains a device identifier.
[0151] In this embodiment of the application, the global data map is a huge data index structure that contains metadata information of data on various terminals. Each metadata node corresponds to a piece of data and records the relevant attributes of the data and the device identifier of the terminal where it is located. In this embodiment, the server performs preliminary screening of all metadata nodes in the global data graph based on constraints. For example, if the constraints specify searching for document data within a specific time range, the server will filter out metadata nodes whose data type is document and whose time is within the specified range, forming a candidate metadata node set with these matching nodes. This can quickly eliminate a large number of nodes that do not meet the conditions, reducing the workload of subsequent calculations.
[0152] In step 4043, the similarity score between the search vector and the semantic vector of each metadata node in the candidate metadata node set is calculated.
[0153] In this embodiment, each candidate metadata node, in addition to containing basic information such as a device identifier, is also associated with a semantic vector. This semantic vector is also a digital representation of the semantic features of the data represented by the metadata node; it extracts key information from the data content and transforms it into vector form using a specific algorithm. In this embodiment of the application, the server can use similarity calculation algorithms (such as cosine similarity, Euclidean distance, etc.) to calculate the similarity score between the search vector and the semantic vector of each candidate metadata node; wherein, the higher the similarity score, the more semantically similar the search content is to the data represented by the metadata node.
[0154] In step 4044, the metadata nodes in the candidate metadata node set are sorted according to the similarity score to generate an ordered search results list.
[0155] In this embodiment, the server sorts all metadata nodes in the candidate metadata node set according to similarity scores. Typically, they are arranged in descending order of similarity score, with higher-scoring nodes appearing first and lower-scoring nodes appearing last. After sorting, an ordered search results list is generated. This list clearly shows the relevance of each candidate metadata node to the search content, facilitating the user's subsequent acquisition of the most relevant data.
[0156] In some embodiments, the metadata nodes in the candidate metadata node set can be sorted in descending order according to the similarity score; based on the preset number of returned results, the top N metadata nodes are selected from the sorted list to form a search results list; where N is an integer greater than zero.
[0157] In this embodiment of the application, the sorting results can also be weighted and adjusted by modifying the time factor during the sorting process.
[0158] In this embodiment, the preset number of returned results N can be flexibly adjusted according to the needs and scenarios of different users. For some simple searches, users may only need a small number of results to meet their needs, such as when searching for weather, returning the first few weather information for different time periods is sufficient; while for complex research topics, users may need more relevant literature for reference, in which case the value of N can be increased to obtain more comprehensive information.
[0159] In this embodiment, the orderly and curated display of results avoids overwhelming users with too much irrelevant information and reduces the frustration caused by information overload. Users can more easily and efficiently find the content they need when browsing search results, thereby improving their satisfaction and user experience throughout the search process.
[0160] In this embodiment, descending order ensures that metadata nodes with high similarity scores are placed first, allowing users to see the data that best matches their search intent and improving information retrieval efficiency.
[0161] In step 4045, the target metadata node is determined from the search results list, and the second terminal is determined based on the device identifier contained in the target metadata node to obtain the positioning result.
[0162] In this embodiment of the application, in the ordered search results list, the server determines the target metadata node according to preset rules (such as selecting the node with the highest similarity score, or filtering out nodes that meet the conditions according to the threshold set by the user). The data represented by this target metadata node is the data that best matches the user's search intent.
[0163] In this embodiment, since each metadata node contains a device identifier, the server can accurately determine the second terminal where the data is located based on the device identifier in the target metadata node. Finally, the device identifier and other information of the second terminal are returned to the first terminal as the location result, completing the location operation of the target data.
[0164] For example, after receiving a response, the intelligent agent of mobile phone A clearly displays a list of results to the user on the interface, such as: 1. Phoenix Project Contract_v3_Draft.docx Location: Zhang San's PC; Modified at: 15:30 yesterday; 2. Phoenix Project Meeting Minutes.pdf Location: Zhang San's PC; Edited 3 days ago.
[0165] In some embodiments, at least one metadata node ranked high in the search results list can be identified as the target metadata node; the device identifier contained in the target metadata node can be extracted to identify the second terminal storing remote target data; a location result can be generated based on the information in the target metadata node; wherein the location result includes at least the device identifier of the second terminal.
[0166] In this embodiment, the search results list is sorted according to the similarity to the search query. Metadata nodes ranked higher indicate a higher relevance to the search content. Identifying these nodes as target metadata nodes ensures that the targeted data is what the user truly needs or best matches their search intent, thus improving the accuracy of the location.
[0167] In a distributed environment, data is stored across multiple terminals. In this embodiment, a global data graph is used to index and manage the data on each terminal device in a unified manner, which can efficiently handle complex cross-device search requests and provide users with a seamless cross-device data search experience.
[0168] In this embodiment, the server can utilize information from the search request to perform an efficient multi-stage composite search within the global data graph. First, constraints are applied for rapid and accurate filtering, significantly narrowing the candidate range from massive metadata nodes. Then, within the filtered candidate set, the similarity between the search vector and the semantic vectors of the candidate nodes is calculated, and the nodes are sorted according to semantic relevance to achieve precise semantic matching. After generating an ordered list of search results, the server can quickly identify the target metadata node without traversing and comparing all nodes, further shortening search time and improving response speed. Furthermore, the combination of search vectors and constraints allows for flexible handling of various search needs, from simple keyword searches to complex semantic searches, accurately matching relevant data and meeting diverse user search needs in a distributed environment.
[0169] As can be seen, in this embodiment of the application, through the above-mentioned refined search and positioning process, it is possible to achieve fast, accurate, and deeply understand user intent cross-device data search without touching the user's original data, and provide accurate positioning results for secure remote operation.
[0170] Figure 5 This is the fifth flowchart of a data processing method provided in some embodiments of this application, applied to a second terminal, such as... Figure 5 As shown, the method may include the following steps: step 501, step 502 and step 503.
[0171] In step 501, the second terminal receives the operation intent data packet sent by the first terminal.
[0172] In this embodiment, the operation intent data packet is constructed by the first terminal and contains rich information to instruct operations on specific remote target data on the second terminal. This information may exist in a specific data structure or encoding form, such as using JSON, XML, or other formats to encapsulate the operation type (e.g., read, write, modify, delete), the identifier of the target data (e.g., data ID, storage path), and possible operation parameters (e.g., the range of data to be read, the specific content of data to be written).
[0173] In this embodiment of the application, the first terminal uses network communication technology, such as wired network (e.g., Ethernet) or wireless network (e.g., Wi-Fi, 4G / 5G, etc.), to send operation intent data packets to the second terminal.
[0174] In step 502, the second terminal parses the operation intent data packet to obtain the operation defined by the operation intent data packet.
[0175] In this embodiment of the application, after receiving the operation intent data packet, the second terminal first performs some preprocessing work, such as checking the integrity of the data packet (through checksum and other methods) to ensure that the data packet is not damaged during transmission.
[0176] In this embodiment, the second terminal parses the operation intent data packet according to a pre-agreed data packet format and parsing rules. For example, if the operation intent data packet is in JSON format, the second terminal uses a corresponding JSON parsing library to extract information such as the operation type, target data identifier, and operation parameters. After parsing, the second terminal understands the specific operation that the first terminal requires to be performed on the remote target data.
[0177] In some embodiments, to ensure the security of the operation intent data packet during transmission, the first terminal encrypts the operation intent data packet before sending it. Upon receiving the operation intent data packet, the second terminal first verifies the digital signature using the first terminal's public key to ensure the integrity and authenticity of the data packet's origin. After successful verification, the second terminal decrypts the data packet using its own private key or a pre-shared key to reconstruct the structured operation intent information, and then parses it to obtain the specific operation instructions. This encryption and verification mechanism effectively prevents data from being tampered with or leaked during transmission, ensuring the security of cross-device operations.
[0178] In some embodiments, to ensure operational security and traceability, the second terminal performs pre-verification on the remote target data based on the parameters in the operation intent data packet before executing the operation. This verification includes, but is not limited to: checking whether the current state of the target data is consistent with expectations, confirming whether the operation permissions are sufficient, and verifying whether the data format conforms to specifications. Only after all pre-verifications are passed will the second terminal execute the corresponding operation. This verification mechanism effectively prevents erroneous operations caused by inconsistent data states and provides traceable operation records for subsequent auditing, thereby enhancing the auditability of the system while ensuring operational security.
[0179] In step 503, the second terminal performs the operation defined in the operation intent data packet on the remote target data and returns the operation execution result to the first terminal.
[0180] In this embodiment, the second terminal performs corresponding operations on the target data locally based on the parsed operation information. For example, if it is a read operation, the second terminal will retrieve the target data from a specified storage location (such as a database, file system, etc.); if it is a write operation, it will write the specified data to the target location. After the operation is completed, the second terminal will package the operation result into a data packet and return it to the first terminal. The result data packet may contain a flag indicating whether the operation was successful, the data content after the operation (if it was a read operation), and error information (if the operation failed). Similarly, the return process will also use network communication technology and may perform data verification to ensure accurate transmission of the result.
[0181] In some embodiments, to prevent file corruption due to accidents such as power outages during modification, the operation is usually atomic. Specifically, the second terminal writes the operation content to a temporary file or temporary area; after confirming successful writing, it submits the changes by renaming or moving the data to replace the original remote target data.
[0182] In some embodiments, the second terminal may also generate and store operation log records locally, wherein the operation log records include at least: operation intent identifier, operation type, operation time, source device identifier and operation status, for subsequent review.
[0183] For example, if the operation is successful, the result of the operation execution is as follows: JSON { "response_to_intent": "op-uuid-67890", / / The original operation ID of the response "status": "SUCCESS", "result": { "new_modification_time": 1753090155 / / Returns the new file modification time. }, "timestamp": 1753090156 }; If the operation fails, the result will be displayed as an example: JSON { "response_to_intent": "op-uuid-67890", "status": "FAILURE", "error": { "code": "E_VERIFICATION_FAILED", Message: "Validation failed: The original value to be modified in the file does not match the expected value." }, "timestamp": 1753090156 }
[0184] In addition, if the operation is successful, a success message (such as a green checkmark) can be displayed on the first terminal; if the operation fails, an error message (such as "Operation failed: the original value in the file does not match") can be displayed on the first terminal to help the user understand the situation.
[0185] As can be seen from the above embodiments, in this embodiment, the first terminal does not need to have in-depth knowledge of the data storage structure and management method of the second terminal. It only needs to construct and send an operation intent data packet to realize the operation of remote target data, which simplifies the operation process of the first terminal and reduces the technical requirements for operators. The entire process does not require the transmission of file content or screen pixel streams, thus improving data security. In addition, the first terminal can flexibly construct operation intent data packets according to actual needs to realize various operations on remote target data, such as real-time reading, batch writing, and conditional modification, enabling the first terminal to adapt to different business scenarios and data management needs.
[0186] The distributed personal cloud system provided in this application embodiment is based on a core architecture of "data content retention locally, metadata logic uploaded to the cloud, and secure command stream interaction," constructing a fully private and distributed personal data processing infrastructure. This system has broad application prospects and can serve as a foundational platform for next-generation personal computing and data collaboration, playing a significant role in multiple fields: Taking highly privacy-sensitive professional fields as an example, in legal compliance scenarios, lawyers can securely search case materials on office computers via mobile terminals, and the system only returns de-identified text fragments of relevant clauses; in healthcare scenarios, doctors can search medical record data on hospital workstations, and the system only returns specific test values, supporting complex operations such as anonymization within a secure domain.
[0187] Taking personal full-domain information management as an example, it can realize unified search and intelligent preview of digital assets (such as photos and documents) across devices, as well as support cross-application task understanding and contextual reminders, and realize information linkage between multiple devices.
[0188] Taking smart device collaboration as an example, this system, acting as a private control hub for smart home devices, enables localized video processing and on-demand segment transmission; and through a global data graph and operational intent commands, it facilitates automated data migration and backup between devices. Through its features of privacy protection, semantic understanding, and efficient collaboration, this system provides a secure and reliable new paradigm for data management across various fields.
[0189] The data processing method provided in this application can be executed by a data processing device. This application uses an example of a data processing device executing the data processing method to illustrate the data processing device provided in this application.
[0190] Figure 6 This is a structural block diagram of a data processing apparatus provided in some embodiments of this application, applied to a first terminal, such as... Figure 6 As shown, the data processing device 600 may include: a generation module 601, a processing module 602, and a sending module 603.
[0191] The generation module 601 is used to generate a local knowledge index based on local data; The processing module 602 is used to perform de-identification processing on the local knowledge index to obtain de-identified metadata; wherein, the metadata includes descriptive information and location information; The sending module 603 is used to upload the metadata to the server; wherein the server constructs a global data graph based on the metadata from at least two terminals. As can be seen, in this embodiment, on the one hand, the terminal performs desensitization processing on the local knowledge index, effectively removing sensitive information that may leak user privacy. Security protection is carried out before the data is uploaded to the server. In this way, even if the metadata is illegally obtained during transmission or storage, attackers cannot obtain the user's sensitive information from it, reducing the risk of data leakage. On the other hand, users do not need to manage data on different terminals separately. They can search the entire data set simply by using the global data graph on the server side, which improves the efficiency of information retrieval in multi-terminal device scenarios.
[0192] Optionally, as an embodiment, the data processing apparatus 600 may further include: The receiving module is used to receive user search requests; The generation module 602 is also used to parse and de-identify the user search request to generate an encrypted search request; The sending module 603 is further configured to send the encrypted search request to the server; wherein, the server obtains the user search request based on the encrypted search request, determines the second terminal where the search data corresponding to the user search request is located based on the global data map, and returns the location result to the first terminal; the location result includes at least the device identifier of the second terminal; The receiving module is also used to receive the positioning result returned by the server.
[0193] Optionally, as an embodiment, the encrypted search request includes constraints, which include at least one of the following: time range, file type, keywords, and hashed entity identifiers.
[0194] Optionally, as an example, the location result may also include at least one of the following: file name, file type, and modification date.
[0195] Optionally, as an embodiment, the positioning result is used to point to remote target data of the second terminal; The receiving module is also used to receive operation instructions for the remote target data; The data processing device 600 may further include: The parsing module is used to obtain the operation intent data packet based on the operation instructions; The sending module 603 is further configured to send the operation intent data packet to the second terminal; wherein the second terminal parses the operation intent data packet, performs the operation defined in the operation intent data packet on the remote target data, and returns the operation execution result to the first terminal; The receiving module is also used to receive the operation execution result returned by the second terminal.
[0196] Optionally, as an embodiment, the operation intent data packet is a JSON object that defines the operation type and operation parameters; wherein, the operation type includes at least one of the following: content update, file deletion, file reading, file renaming; and the operation parameters include at least one of the following: location information of remote target data, search context, original value to be verified, and new value to be set.
[0197] Optionally, as an embodiment, the data processing apparatus 600 may further include: The encryption module is used to encrypt the operation intent data packet using the public key or pre-shared symmetric key of the second terminal, and to digitally sign the encrypted data packet using the private key of the first terminal.
[0198] Optionally, as an embodiment, the operation execution result is encrypted and signed status confirmation information; wherein, the status confirmation information includes a status code indicating whether the operation was successful or failed, as well as necessary error information, but does not include the original content of the remote target data.
[0199] Optionally, as an embodiment, the generation module 601 is specifically used to perform entity recognition, keyword extraction, and semantic vectorization processing on local data through a natural language processing model to generate a local knowledge index containing semantic information.
[0200] Optionally, as an embodiment, the processing module 602 is configured to perform at least one of the following processes: perform an irreversible hash operation on the identified named entities to generate a hash value; identify and filter out sensitive information with a preset format; and replace specific numerical information with abstract type identifiers or range quantization values.
[0201] Optionally, as an embodiment, the location information includes: a data pointer; wherein, the description information includes at least one of the following: basic structure information and content-derived semantic information; the basic structure information is used to describe the basic physical and structural properties of the data; the content-derived semantic information is used to describe the content of the data, including at least one of the following: semantic vector embedding, hashed named entity list, and structural summary.
[0202] The data processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0203] The data processing device in this application embodiment can be a device with an operating system. The operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system.
[0204] The data processing device provided in this application embodiment can achieve the above-mentioned... Figures 1 to 3 To avoid repetition, the various processes implemented in any of the method embodiments shown will not be described again here.
[0205] Optionally, such as Figure 7 As shown, this application embodiment also provides an electronic device 700, including a processor 701 and a memory 702. The memory 702 stores a program or instructions that can run on the processor 701. When the program or instructions are executed by the processor 701, they implement the various steps of the above-described data processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0206] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0207] Figure 8 This is a schematic diagram of the hardware structure of an electronic device that implements various embodiments of this application. The electronic device 800 includes, but is not limited to, components such as: a radio frequency unit 801, a network module 802, an audio output unit 803, an input unit 804, a sensor 805, a display unit 806, a user input unit 807, an interface unit 808, a memory 809, and a processor 810.
[0208] In some embodiments, the electronic device 800 is a first terminal, and the processor 810 is configured to generate a local knowledge index based on local data; perform desensitization processing on the local knowledge index to obtain desensitized metadata; wherein the metadata includes descriptive information and location information; and upload the metadata to a server; wherein the server constructs a global data map based on metadata from at least two terminals.
[0209] Optionally, as an embodiment, the processor 810 is further configured to receive a user search request; parse and de-identify the user search request to generate an encrypted search request; send the encrypted search request to the server; wherein the server obtains the user search request based on the encrypted search request, determines the second terminal where the search data corresponding to the user search request is located based on the global data map, and returns the location result to the first terminal; the location result includes at least the device identifier of the second terminal; and receives the location result returned by the server.
[0210] Optionally, as an embodiment, the encrypted search request includes constraints, which include at least one of the following: time range, file type, keywords, and hashed entity identifiers.
[0211] Optionally, as an example, the location result may also include at least one of the following: file name, file type, and modification date.
[0212] Optionally, as an embodiment, the positioning result is used to point to remote target data of the second terminal; The processor 810 is further configured to receive an operation instruction for the remote target data; obtain an operation intent data packet according to the operation instruction; send the operation intent data packet to the second terminal; wherein the second terminal parses the operation intent data packet, performs the operation defined in the operation intent data packet on the remote target data, and returns the operation execution result to the first terminal; and receives the operation execution result returned by the second terminal.
[0213] Optionally, as an embodiment, the operation intent data packet is a JSON object that defines the operation type and operation parameters; wherein, the operation type includes at least one of the following: content update, file deletion, file reading, file renaming; and the operation parameters include at least one of the following: location information of remote target data, search context, original value to be verified, and new value to be set.
[0214] Optionally, as an embodiment, the processor 810 is further configured to encrypt the operation intent data packet using the public key or pre-shared symmetric key of the second terminal, and to digitally sign the encrypted data packet using the private key of the first terminal.
[0215] Optionally, as an embodiment, the operation execution result is encrypted and signed status confirmation information; wherein, the status confirmation information includes a status code indicating whether the operation was successful or failed, as well as necessary error information, but does not include the original content of the remote target data.
[0216] Optionally, as an embodiment, the processor 810 is specifically used to perform entity recognition, keyword extraction, and semantic vectorization processing on local data through a natural language processing model to generate a local knowledge index containing semantic information.
[0217] Optionally, as an embodiment, the processor 810 is specifically configured to perform at least one of the following: perform an irreversible hash operation on the identified named entities to generate a hash value; identify and filter out sensitive information with a preset format; and replace specific numerical information with an abstract type identifier or range quantization value.
[0218] Optionally, as an embodiment, the location information includes: a data pointer; wherein, the description information includes at least one of the following: basic structure information and content-derived semantic information; the basic structure information is used to describe the basic physical and structural properties of the data; the content-derived semantic information is used to describe the content of the data, including at least one of the following: semantic vector embedding, hashed named entity list, and structural summary.
[0219] It should be understood that, in this embodiment, the input unit 804 may include a graphics processing unit (GPU) 8041 and a microphone 8042. The GPU 8041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 806 may include a display panel 8061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 807 includes at least one of a touch panel 8071 and other input devices 8072. The touch panel 8071 is also called a touch screen. The touch panel 8071 may include two parts: a touch detection device and a touch controller. Other input devices 8072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0220] The memory 809 can be used to store software programs and various data. The memory 809 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 809 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (Synchlink DRAM, SLDRAM), and direct memory bus RAM (DRRAM). The memory 809 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0221] Processor 810 may include one or more processing units; optionally, processor 810 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 810.
[0222] This application also provides a readable storage medium storing a program or instructions. When executed by a processor, the program or instructions implement the various processes of the above-described data processing method embodiments and achieve the same technical effects. To avoid repetition, further details are omitted here. The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.
[0223] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described data processing method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here. It should be understood that the chip mentioned in this application embodiment can also be called a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0224] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the data processing method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0225] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0226] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0227] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A data processing method, characterized in that, Applied to a first terminal, the method includes: Generate a local knowledge index based on local data; The local knowledge index is anonymized to obtain anonymized metadata; wherein the metadata includes descriptive information and location information; The metadata is uploaded to a server; wherein the server constructs a global data graph based on metadata from at least two terminals.
2. The method according to claim 1, characterized in that, The method further includes: Receive user search requests; The user search request is parsed and de-identified to generate an encrypted search request; The encrypted search request is sent to the server; wherein, the server obtains the user search request based on the encrypted search request, determines the second terminal where the search data corresponding to the user search request is located based on the global data map, and returns the location result to the first terminal; the location result includes at least the device identifier of the second terminal; Receive the location result returned by the server.
3. The method according to claim 2, characterized in that, The encrypted search request includes constraints, which include at least one of the following: time range, file type, keywords, and hashed entity identifiers.
4. The method according to claim 2, characterized in that, The location results also include at least one of the following: file name, file type, and modification date.
5. The method according to claim 2, characterized in that, The location result is used to point to remote target data of the second terminal; after receiving the location result returned by the server, the method further includes: Receive operation instructions for the remote target data; The operation intent data packet is obtained according to the operation instructions; The operation intent data packet is sent to the second terminal; wherein, the second terminal parses the operation intent data packet, performs the operation defined in the operation intent data packet on the remote target data, and returns the operation execution result to the first terminal; Receive the operation execution result returned by the second terminal.
6. The method according to claim 5, characterized in that, The operation intent data packet is a JSON object that defines the operation type and operation parameters; The operation types include at least one of the following: content update, file deletion, file reading, and file renaming; The operation parameters include at least one of the following: location information of the remote target data, search context, original value to be verified, and new value to be set.
7. The method according to claim 5, characterized in that, Before sending the operation intent data packet to the second terminal device, the method further includes: The operation intent data packet is encrypted using the public key or pre-shared symmetric key of the second terminal, and the encrypted data packet is digitally signed using the private key of the first terminal.
8. The method according to claim 5, characterized in that, The result of the operation is an encrypted and signed status confirmation message; The status confirmation information includes a status code indicating whether the operation was successful or failed, as well as necessary error information, but does not include the original content of the remote target data.
9. The method according to claim 1, characterized in that, The process of generating a local knowledge index based on local data includes: By using a natural language processing model, local data is processed for entity recognition, keyword extraction, and semantic vectorization to generate a local knowledge index containing semantic information.
10. The method according to claim 1, characterized in that, The desensitization process for the local knowledge index includes at least one of the following: Perform an irreversible hash operation on the identified named entities to generate hash values; Identify and filter out sensitive information with a preset format; Replace specific numerical information with abstract type identifiers or range quantization values.
11. The method according to claim 1, characterized in that, The location information includes: a data pointer; The descriptive information includes at least one of the following: basic structural information and content-derived semantic information; The infrastructure information is used to describe the basic physical and structural properties of the data; The content-derived semantic information is used to describe the content of the data, including at least one of the following: semantic vector embedding, hashed named entity list, and structural summary.
12. A data processing apparatus, characterized in that, Applied to a first terminal, the device includes: The generation module is used to generate a local knowledge index based on local data. The processing module is used to perform de-identification processing on the local knowledge index to obtain de-identified metadata; wherein, the metadata includes description information and location information; A sending module is used to upload the metadata to a server; wherein the server constructs a global data graph based on metadata from at least two terminals.
13. The apparatus according to claim 12, characterized in that, The device further includes: The receiving module is used to receive user search requests; The generation module is also used to parse and de-identify the user search request to generate an encrypted search request; The sending module is further configured to send the encrypted search request to the server; wherein, the server obtains the user search request based on the encrypted search request, determines the second terminal where the search data corresponding to the user search request is located based on the global data map, and returns the location result to the first terminal; the location result includes at least the device identifier of the second terminal; The receiving module is also used to receive the positioning result returned by the server.
14. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing programs or instructions that can run on the processor, the programs or instructions being executed by the processor to implement the steps of the data processing method applied to a first terminal as described in any one of claims 1 to 11.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the data processing method applied to a first terminal as described in any one of claims 1 to 11.