A method and system for managing and storing business candidate information
By constructing an irreversible semantic representation model and a multi-perturbation matrix strategy, the problem of sensitive information leakage in the candidate information management system is solved, and secure and controllable data access and cross-organizational information sharing are achieved.
Patent Information
- Application Number
- CN202510885151.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-30
AI Technical Summary
In the existing technology, while ensuring semantic responsiveness, the candidate information management system lacks structural irreversibility, disturbance resistance and access controllability, resulting in the risk of sensitive information leakage and privacy infringement.
By constructing an initial representation vector group and adopting irreversible semantic modeling and multi-perturbation matrix strategies, the field information of candidate personnel is converted into an irreversible semantic representation model, and stored separately from the field information data. Access is controlled only through index paths, achieving data response rather than exposure.
It improves the management and control security of sensitive information and the controllability of model calls, ensures the isolation and security of field information, prevents data leakage, and is suitable for cross-organizational talent model sharing and analysis.
Smart Images

Figure CN120372669B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage, and in particular to a method and system for managing and storing business candidate information. Background Art
[0002] In existing talent information management systems, candidate field information is typically stored centrally in a structured format and retrieved and analyzed through field-level indexing. However, this type of field information often contains sensitive content such as name, contact information, work history, and competency tags. Once illegally accessed or desensitization fails, it can easily lead to data leakage, privacy violations, and organizational security risks. To improve access efficiency and intelligence, some systems have introduced vectorized representation mechanisms, but most solutions still retain field-level semantic boundaries or have field restore capabilities, and cannot technically block the mapping link from semantic expression to field content.
[0003] For example, a Chinese patent application with publication number CN120012079A provides a text recognition method based on a language model. The method includes obtaining a text data stream to be detected within a target application, including user input, interface display, and network transmission text; preprocessing the text data stream to generate a multi-dimensional semantic vector sequence; modulating the semantic vector sequence through a dynamic weight allocation module to generate a modulation feature matrix; inputting a logic judgment engine to perform multi-level condition combination analysis to generate a risk probability distribution; demodulating the risk probability distribution to generate an interpretable risk feature vector; performing cross-validation based on the risk feature vector and a preset threshold set to generate a final risk determination result; and triggering a multi-layer protection mechanism, including content replacement, session interruption, and security alerts, when a high risk is determined. This invention can improve the accuracy and interpretability of the system, ensuring the security of the target application and user experience.
[0004] The above existing technologies all have the problems raised by this background technology: there is a lack of a candidate personnel information management method that has structural irreversibility, disturbance resistance and access controllability while ensuring semantic responsiveness. In order to solve the above problems, this application designs a business candidate personnel information management method and system. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to address the shortcomings of the existing technology and provide a business candidate personnel information management and storage method and system. The method obtains multiple field information data of business candidate personnel and constructs an initial representation vector group. Irreversible semantic modeling is performed according to a predefined perturbation strategy to generate a semantic embedding representation and encapsulate it into a semantic representation model. The perturbation strategy includes dividing the feature subspace, setting perturbation matrices with different tuning rules, and asymmetric coupling perturbation based on cross-feedback. The semantic model is stored separately from the field information data, and the fields are only accessed through controlled index paths, achieving data responsiveness rather than exposure. This method improves the management and control security of sensitive information and the controllability of model calls.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A method for managing and storing business candidate information, the method comprising:
[0008] Acquiring multiple field information data of the business candidate, performing feature extraction and encoding on the multiple field information data, and constructing an initial representation vector group;
[0009] Performing irreversible semantic modeling on the initial representation vector group according to a predefined perturbation strategy, wherein the perturbation strategy includes:
[0010] Set the initial representation vector group and reconstruct the error distribution, and obtain the feature subspace based on the main dimension cluster of the error;
[0011] Initialize the perturbation matrix for each characteristic subspace, and the tuning rules and tuning amounts of each perturbation matrix are different;
[0012] The perturbation matrix acts on the corresponding feature subspace and generates a cross-feedback vector according to the perturbation response of the feature subspace, generates a secondary perturbation response according to the cross-feedback vector to update the corresponding feature subspace, generates a semantic embedding representation, and generates a semantic representation model according to the semantic embedding representation;
[0013] The semantic representation model is stored in a preset information database.
[0014] The disruption strategies include:
[0015] Perform the first round of random dimensionality reduction on the initial representation vector group to generate the first round of dimensionality reduction embedding vector;
[0016] Performing an inverse reconstruction calculation on the first round of dimensionality reduction embedding vectors to obtain a reconstruction error distribution feature between the vectors and the initial representation vector group;
[0017] Initializing at least two perturbation matrices according to the reconstruction error distribution characteristics, respectively acting on different feature subspaces of the embedding vector obtained by the first round of dimensionality reduction;
[0018] Perturbing the corresponding feature subspace according to the perturbation matrix, and generating a local perturbation expression according to the perturbation response, wherein the perturbation is also cross-acted based on the perturbation response of the feature subspace corresponding to other perturbation matrices;
[0019] The local perturbation expressions are nested and combined to generate a semantic embedding representation.
[0020] Initializing at least two disturbance matrices according to the reconstruction error distribution characteristics, including:
[0021] Calculating the error intensity distribution of the reconstruction error distribution feature in different feature dimensions;
[0022] Dividing the first round of dimensionality reduction embedding vector into at least two feature subspaces according to the error intensity distribution, wherein each feature subspace corresponds to a main dimension cluster of error intensity;
[0023] Initializing a main perturbation kernel of a perturbation matrix according to the direction vector of the main dimension cluster, wherein the main perturbation kernel includes a perturbation basis vector constructed along the direction vector and a Gaussian perturbation term;
[0024] According to the statistical characteristics of the error in the characteristic subspace and the structural characteristics of the main perturbation kernel, the tuning rule and tuning amount of the perturbation matrix are determined.
[0025] The tuning rules include a disturbance direction update method, an amplitude gain strategy, and a disturbance activation condition, and the tuning variables include a disturbance intensity range, an update step, and a disturbance probability threshold.
[0026] Perturbing the corresponding feature subspace according to the perturbation matrix includes:
[0027] Performing a perturbation operation on the feature embedding vectors in the respective feature subspaces according to the perturbation matrix to generate a preliminary perturbation response, wherein the perturbation operation includes feature shifting and injecting random Gaussian noise;
[0028] For any perturbation matrix, when performing the perturbation operation, a cross-feedback vector is generated according to the preliminary perturbation response of the characteristic subspace corresponding to the other perturbation matrix;
[0029] The disturbance direction of the current disturbance matrix is adjusted according to the cross-feedback vector to generate a secondary disturbance response.
[0030] Adjusting the disturbance direction of the current disturbance matrix according to the cross feedback vector includes:
[0031] Calculating the angle between the cross feedback vector and the feature subspace direction vector;
[0032] The disturbance direction of the current disturbance matrix is rotated around the cross feedback vector by the angle to generate an updated disturbance direction.
[0033] When an external access request is received, the information storage method further includes:
[0034] Parsing the external access request;
[0035] matching semantic response dimensions from the semantic representation model according to the parsing results;
[0036] Generate a response result according to the semantic response dimension, wherein the response result queries the field information data according to the index path;
[0037] After the response is completed, the semantic representation model is encapsulated according to the access control configuration.
[0038] Matching the semantic response dimension from the semantic representation model according to the parsing result includes:
[0039] Build access intent vector;
[0040] Matching the access intention vector with the semantic embedding representation in the semantic representation model, wherein the matching includes semantic similarity calculation and structural path intersection;
[0041] Filter semantic nodes whose semantic similarity is greater than a preset threshold or whose path overlap is greater than a preset condition from the matching results;
[0042] The index dimension set covered by the semantic node is used as the semantic response dimension.
[0043] The access control configuration includes an access entropy adjustment policy, and the access entropy adjustment policy includes:
[0044] Calculating an access entropy value, wherein the access entropy value is calculated based on the nesting level depth, the number of response dimensions, the structural path density, and the coverage ratio of the original field index path of the semantic response dimension hit by the access request in the semantic representation model;
[0045] An encapsulation strategy is selected according to the access entropy value, wherein the encapsulation strategy includes structure compression encapsulation, summary encapsulation and scrambled encapsulation.
[0046] A business candidate information management and storage system, the system comprising:
[0047] The modeling module is used to obtain field information data of business candidates and build an irreversible vectorized representation model containing a semantic nested structure;
[0048] The perturbation module is used to perform random dimensionality reduction, error analysis, and cross-control of multiple perturbation matrices on the initial representation vector to generate an irreversible semantic embedding representation;
[0049] A storage module is used to store the semantic representation model and field information data separately and establish an index path connection structure;
[0050] The access module is used to parse external access requests, match semantic response dimensions, and generate response results through index paths;
[0051] The encapsulation module is used to calculate the access entropy value according to the access entropy adjustment strategy, and select the corresponding encapsulation method to perform structural compression, summary and scrambling operations on the semantic representation model.
[0052] Compared with the prior art, the present invention has the following beneficial effects:
[0053] The present invention converts candidate field information into an irreversible semantic representation model, introduces multiple perturbation matrices to act on different feature subspaces, and combines tuning rule differences with a cross-feedback mechanism to achieve deep perturbation of the mapping structure between field information and semantic expression, effectively interrupting the reversible path of the field. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0055] Figure 1 This is a schematic diagram of an exemplary application scenario of an embodiment of the present invention;
[0056] Figure 2 This is a flow chart of a method for managing and storing business candidate information according to an embodiment of the present invention;
[0057] Figure 3 The figure is a flow chart of a data conversion method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0058] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0059] References to "embodiments" herein mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it refer to independent or alternative embodiments that are mutually exclusive of other embodiments. It will be understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0060] See also Figure 1 , which is a schematic diagram of an exemplary application scenario provided in an embodiment of the present application.
[0061] like Figure 1 As shown in the figure, the application scenario involves sharing talent models among multiple subsidiaries within a group, including Subsidiary A, Subsidiary B, and Subsidiary C. All three subsidiaries have local candidate field data. To prevent the direct upload of sensitive fields, each subsidiary converts the candidate field data into a semantic representation model using irreversible modeling.
[0062] Figure 1 It shows that subsidiary A generates semantic representation model A, subsidiary B generates semantic representation model B, and subsidiary C generates semantic representation model C.
[0063] In one example, the above semantic representation models are all irreducible vector expression structures.
[0064] Figure 1 This shows that after each subsidiary completes semantic modeling locally, it uploads its generated semantic representation model to the headquarters platform talent pool. The headquarters platform centrally stores the semantic representation models and uses them to respond to access requests from within the platform or authorized systems.
[0065] Figure 1 The diagram shows that when the headquarters platform receives an access request, it matches the request content to the corresponding semantic representation model, generates a semantic response based on the matching results, and returns the response result to the accessing party. Throughout this process, candidate information is not transmitted, stored, or displayed on the headquarters platform. The access request is responded to solely through the irreversible semantic representation model, ensuring the isolation and security of field information.
[0066] In one example, the semantic representation model can be used by the headquarters platform to perform talent capability modeling and job matching across subsidiaries. By defining a unified semantic nested structure, semantic models uploaded by different subsidiaries are structurally comparable. The platform can calculate semantic similarity between candidates and positions based on semantic dimensions, which can be used for business needs such as job recommendations, echelon building, and organizational capability prediction.
[0067] In one example, the platform can generate a candidate capability map based on an aggregated semantic representation model. This map does not rely on field-level attributes, but only relies on semantic structure and feature transformation results. It is suitable for talent distribution analysis, capability gap identification and organizational strategy assessment, and has cross-system, cross-language and cross-organizational adaptability.
[0068] On the other hand, the irreversible modeling method can be deployed in a private isolated computing environment, allowing only local use of field information data to complete semantic vector generation, and destroying the intermediate field cache immediately after the model is generated, without generating the link risk of field-level data leakage, and meeting the data compliance and cross-organizational model collaboration security requirements in special scenarios.
[0069] Next, a method for managing and storing business candidate information provided by an embodiment of the present application will be introduced with reference to the accompanying drawings.
[0070] See also Figure 2 , this figure is a flow chart of a method for managing and storing business candidate information provided in an embodiment of the present application. Figure 2 The method shown can be applied within the group, on government-enterprise platforms, or in multi-source talent collaboration systems to perform semantic modeling and access protection on candidate field data. Figure 2 The method shown includes the following steps S1-S3, and the specific steps are as follows:
[0071] S1: Obtain multiple fields of information data of business candidates;
[0072] In this step, the field information data may include data content containing field information extracted from the company's internal database, human resources system or third-party candidate interface. The fields include but are not limited to name, education, job experience, ability tags and other structured or semi-structured information.
[0073] S2: Convert field information data into a semantic representation model;
[0074] In this step, based on the predefined semantic modeling mechanism, feature extraction and encoding of field information data are performed to construct an initial representation vector group;
[0075] Then, an irreversible modeling strategy is executed, which includes random dimensionality reduction, reconstruction error calculation, perturbation matrix initialization and cross-feedback perturbation, to generate an irreversible vectorized expression of the semantic nested structure, where the irreversible vectorized expression is the semantic embedding representation.
[0076] The semantic embedding representation is encapsulated into a semantic representation model, which does not contain directly identifiable field semantic boundaries and is used to drive semantic layer data interaction when responding to access requests later.
[0077] S3: Separately stores the semantic representation model and field information data, and establishes an access index path;
[0078] In this step, the semantic representation model is bound to the access control configuration and stored separately from the field information data. The field information data is placed in a temporary access cache and is only accessed through the index path based on the semantic response dimension when the access control conditions are met. This ensures that the field information data is not leaked or reconstructed during the entire modeling and calling process.
[0079] Although semantic modeling and feature vector expression have been widely used in existing technologies for tasks such as talent recommendation, resume screening, and candidate capability matching, the current common practice is still based on reducible structures or weak semantic isolation mechanisms. Even after desensitization, the vector structure itself can still be used to restore part or all of the original field information through model reverse reasoning or field distribution backpropagation. The typical feature of this type of technical solution is that its vector expression, while preserving the field semantic boundaries, field context structure, and encoding path, fails to effectively interrupt the correspondence between the semantic structure and the original field. Therefore, during the model call process, attackers can access the semantic interface at high frequency, construct a dimensional frequency map, a semantic similarity map, or a gradient reverse path, and gradually restore the candidate's sensitive fields, resulting in privacy leakage or resource competition risks.
[0080] In this embodiment, by constructing a random initial dimensionality reduction space and introducing a multi-perturbation matrix mechanism driven by error feedback, the system performs multiple rounds of asymmetric and structurally decoupled dimensionality reduction operations on the original field vector group, so that the semantic embedding model finally generated no longer retains the semantic boundaries between fields in structure, and the vector expression itself no longer has a reversible mapping path. Compared with the traditional embedding method, the present application scheme breaks the stable mapping relationship between fields and semantic representations while maintaining the responsiveness of the model. It should be noted that the present application method uses a cross-perturbation strategy between multiple matrices, so that each perturbation response is not only independent of the field source, but also coupled with the subspace feedback of other perturbation matrices, thereby constructing an expression structure that is only responsive in local dimensions but not globally inferable. This structure naturally has the ability to resist reconstruction. Even if the attacker has some field labels and the semantic response results of the model, it is difficult to reconstruct the original field vector or identify the sensitive entity corresponding to the semantic expression unit.
[0081] For example, in a talent sharing scenario within a group, multiple subsidiaries each have the original field information of independent candidates. Traditional solutions may desensitize the fields and generate embedding vectors before uploading them to the headquarters platform. However, since these vectors still retain the field semantic path in structure, there is a risk of restoration. In the implementation method of the present application, each subsidiary performs multi-matrix perturbation modeling on the field information locally and only uploads the irreversible semantic model to the headquarters. When the headquarters platform executes a matching request for a certain position, it can only obtain the semantic response result, and the response path is controlled by access entropy to ensure that even the headquarters platform cannot restore the candidate fields. In this way, the platform can complete semantic layer operations such as talent recommendation and capability analysis without involving any field leakage issues.
[0082] In this embodiment, since field information data typically have different semantic granularity, numerical types, and statistical distribution characteristics, direct vectorization may result in insufficient semantic coupling or imbalanced embedding space dimensions, thereby affecting the quality of subsequent irreversible modeling. Therefore, before semantic modeling, this embodiment preferably performs structured preprocessing on the original field information data, specifically including field screening, type normalization, and multi-scale semantic reconstruction. Among them, field screening is used to eliminate field information with weak relevance or poor stability to the modeling task, such as extremely sparse labels, low-relevance evaluation items, etc., to reduce the interference of redundant dimensions on the construction of the perturbation matrix. Type normalization includes performing normalization transformation on numerical fields, distributed encoding or multi-hot vector expansion on enumeration fields, and using word embedding, context aggregation, or label mapping to uniformly represent text fields, thereby constructing an initial representation vector group. To improve semantic adversarial capabilities, multi-scale semantic construction mechanisms can also be introduced on some fields, such as through field segmentation, nested label construction, or dynamic context encoding to generate hierarchical field representations to enhance the semantic elasticity retained after dimensionality reduction perturbation.
[0083] See also Figure 3 , which is a flow chart of a data conversion method provided in an embodiment of the present application. Figure 3 The method shown can be applied to step S2 of the aforementioned method, and the specific steps are as follows:
[0084] S2.1: Extract and encode the features of the plurality of field information data to construct an initial representation vector group;
[0085] Specifically, due to significant differences in structural type, numerical distribution, and semantic boundaries between raw field information, even modeling the fields after preprocessing can still easily lead to feature redundancy, dimensional noise, or incomplete semantic coupling, which in turn affects the effectiveness of subsequent disruption strategies. This embodiment does not directly use traditional language model embedding. Instead, it generates a set of semantic features at a controllable scale through key phrase screening, stem mapping, and semantic label construction, and embeds them into the structured representation stream using a windowed splicing method.
[0086] In this embodiment, the construction of the initial representation vector group also includes a field-level attention filtering mechanism. Based on the average importance or variance distribution of candidate fields in historical training samples, fields with significant weight contributions are dynamically selected for encoding, reducing the interference of non-critical fields on the overall semantic modeling. The vector group is ultimately represented as a two-dimensional tensor, with each row corresponding to a feature segment or semantic unit of a field, with uniform length and semantic level classification identifiers.
[0087] S2.2: Perform irreversible semantic modeling on the initial representation vector group according to a predefined perturbation strategy to generate a semantic embedding representation;
[0088] The perturbation strategy is used to perturb the mapping structure between the initial representation vector group and the feature space;
[0089] The goal of this step is to construct a semantic structure representation that retains expressive power under the response task but cannot be restored to a field-level representation through modeling or computation. The perturbation strategy, based on mechanisms such as random dimensionality reduction, reconstruction error feedback, and cross-coupling of multiple perturbation matrices, weakens the original semantic boundaries and interrupts the field dimension path, forming an irreversible vector representation. During the perturbation process, the semantic embedding representation is limited to a low-dimensional, compressed, and hybrid representation format. The field-level structure and label traces are not retained, but the semantic density and response gradient based on task optimization are still preserved.
[0090] The specific steps of step S2.2 of the above method are as follows:
[0091] S2.2.1: Perform a first round of random dimensionality reduction on the initial representation vector group to generate the first round of dimensionality reduction embedding vectors;
[0092] Specifically, to disrupt the continuous expression path of the original field information in the feature space, the first step in building an irreversible semantic model is to introduce a randomized dimensionality reduction mechanism to map the high-dimensional field representation into a low-dimensional semantic embedding space. The core goal is to eliminate the semantic boundaries and distributional inertia of the fields during the original encoding process, forcing the feature representation to undergo unpredictable projective transformations. By reducing the dimensionality, the original distance structure and category distribution between fields are compressed or distorted, thereby reducing the possibility of restoration.
[0093] In this embodiment, a dimensionally matched random Gaussian matrix is first generated as the dimensionality reduction transformation matrix. This matrix is not trained on the model and is initialized only once when constructing the semantic embedding. Its row vectors are normalized to form a low-rank projection basis. Subsequently, matrix multiplication is used to project the original high-dimensional initial representation vector group into a low-dimensional space to generate the first round of dimensionality reduction embedding vectors. This dimensionality reduction result does not retain field-level semantic boundaries or rely on a preset dictionary structure, thus possessing initial irreversibility while retaining the overall coarse-grained semantic distribution trend.
[0094] As an example, in order to avoid the weakening of semantic expression ability caused by information loss introduced by dimensionality reduction, the dimension of the projection matrix sets the balance between the information entropy of the reference field and the structural redundancy, and is generally controlled between 1 / 4 and 1 / 6 of the original dimension, thereby ensuring that there is still sufficient semantic carrying capacity in the compressed space.
[0095] S2.2.2: Perform inverse reconstruction calculation on the first round of dimensionality reduction embedding vectors to obtain a distribution characteristic of the reconstruction error between the vectors and the initial representation vector group;
[0096] Specifically, this step aims to assess the degree of semantic distortion introduced by random dimensionality reduction and, based on this, identify highly sensitive regions to inform differentiated strategies for constructing the subsequent perturbation matrix. Directly performing a homogeneous perturbation on all reduced dimensionality vectors would make it difficult to control the risk of semantic loss and would not be able to focus on perturbations of sensitive dimensions.
[0097] In this embodiment, the first round of dimensionality reduction embedding vectors is reconstructed using a pseudo-inverse matrix. Linear backprojection of the embedding vectors is performed using the pseudo-inverse of the previous random Gaussian matrix to recover an approximate original representation. Subsequently, the Euclidean distance between the reconstructed vector and the true initial representation vector is calculated in each dimension to obtain the reconstruction error value for each dimension, ultimately constructing a complete set of error intensity vectors. These error vectors are normalized and smoothed using a sliding window to determine the central tendency of error information for each semantic dimension.
[0098] S2.2.3: Initialize at least two perturbation matrices based on the reconstruction error distribution characteristics, wherein the perturbation matrices have different tuning rules and tuning amounts, and act on different characteristic subspaces of the embedding vectors in the first round of dimensionality reduction;
[0099] Specifically, based on the error-dominant dimension clusters identified in the previous stage, the embedding vector from the first round of dimensionality reduction is divided into multiple feature subspaces, each corresponding to a dimension cluster. The core purpose of constructing multiple perturbation matrices is to employ differentiated perturbation strategies for irreversible processing of different feature subspaces, thereby increasing the perturbation complexity of the overall structure in multiple directions and avoiding the formation of predictable perturbation trajectories under a single perturbation direction.
[0100] In this embodiment, an independent perturbation matrix is constructed for each characteristic subspace. The perturbation matrix contains a set of main perturbation kernels initialized along the dominant direction of the error, and its structure consists of the main direction vector of the error cluster and a set of Gaussian perturbation terms. The perturbation kernel ensures the independence of each perturbation matrix in the perturbation strategy by controlling the perturbation norm, perturbation sparsity, and the nonlinear coupling relationship between the covariance matrix and other perturbation kernels. Each perturbation matrix also sets independent tuning rules and tuning amounts, including the update strategy of the perturbation direction, the triggering conditions for perturbation activation, the amplitude control of the perturbation gain, the dynamic adjustment method of the perturbation probability threshold, etc.
[0101] In an example, the specific steps for initializing the perturbation matrix are as follows:
[0102] S2.2.3.1: Calculate the error intensity distribution of the reconstruction error distribution feature at different feature dimensions;
[0103] S2.2.3.2: Divide the first round of dimensionality reduction embedding vector into at least two feature subspaces according to the error intensity distribution, where each feature subspace corresponds to a main dimension cluster of error intensity;
[0104] S2.2.3.3: Initialize a main perturbation kernel of the perturbation matrix according to the direction vector of the main dimension cluster, wherein the main perturbation kernel includes a perturbation basis vector constructed along the direction vector and a Gaussian perturbation term;
[0105] S2.2.3.4: Determine the tuning rules and tuning amount of the perturbation matrix based on the statistical characteristics of the error in the characteristic subspace and the structural characteristics of the main perturbation kernel.
[0106] In this embodiment, the perturbation matrix initialization step is used to construct multiple heterogeneous perturbation matrices based on the specific characteristic structure of the reconstruction error distribution after the initial representation vector group completes the first round of random dimensionality reduction, thereby achieving the goal of structural perturbation in the semantic embedding representation. The perturbation matrices are not generated using a unified template. Instead, they are constructed by spatially decomposing the reconstruction deviation trends of the reduced dimensionality vectors in different dimensions and then partitioning them. Differentiated response mechanisms are formed within each subspace, ultimately achieving a nested, coupled, and irreversible perturbation.
[0107] As an example, in S2.2.3.1, the distribution of the reconstruction error between the first round of dimensionality reduction embedding vector and the original initial representation vector group is calculated. The original intention of the design of this step is to identify which feature dimensions or dimension groups have a large degree of semantic information loss during the dimensionality reduction and compression process, so as to provide a distribution basis for subsequent perturbation operations. Although the dimensionality reduction operation can be regarded as an information-preserving transformation in theory, in the actual modeling process, due to factors such as redundancy, collinearity or uneven semantic local density in high-dimensional vectors, the dimensionality reduction process is often accompanied by structural decoupling of some highly sensitive dimensions, making the semantic reconstruction accuracy of some areas significantly lower than that of other dimensions. Therefore, making these dimensional differences explicit and using them to distinguish perturbation strategies has become one of the key points to improve perturbation accuracy and irreversibility.
[0108] In this embodiment, the reconstruction error is calculated using a pseudo-inverse reprojection technique. This technique approximates the original representation vectors through inverse mapping without introducing a learnable mapping model, and then obtains the error value for each dimension through vector differentiation. To eliminate the influence of local noise on the error clustering results, this step also introduces sliding window smoothing, interval truncation, and local statistical fitting to ensure the continuity and cluster operability of the resulting error distribution curve. The resulting error distribution vector serves as the basis for subspace partitioning and also provides upper and lower boundary conditions for perturbation intensity adjustment.
[0109] As an example, in S2.2.3.2, based on the obtained reconstruction error distribution results, the first round of dimensionality reduction embedding vector is divided into several feature subspaces according to the error-dominated dimension. This division is not a simple equal slicing operation, but a dynamic clustering division based on the error density distribution, so that the dimensions in each feature subspace have a high degree of similarity in reconstruction performance, or have similar semantic sensitivity indicators. Since the dimension mapping relationship in the dimensionality reduction process often presents non-uniform characteristics of local compression and global sparsity, directly using a unified perturbation strategy to act on the entire embedding space may cause some important information to be omitted or disturbed, or some low-noise areas to be mistakenly disturbed, causing semantic ambiguity.
[0110] This embodiment does not use a unified perturbation matrix to act on the entire embedding space. Instead, it divides the embedding space into at least two or more feature subspaces and designs its own perturbation response logic within each subspace. This subspace division method not only improves the perturbation's targeting, but also avoids problems such as mutual interference and signal shielding between different regions during the perturbation process. It helps to build a local perturbation structure with stronger controllability and more independent action paths, thereby improving the anti-reduction ability of the overall embedding expression model.
[0111] As an example, in S2.2.3.3, a perturbation matrix is initialized for each characteristic subspace, and its core is composed of the main perturbation kernel. The construction logic of the main perturbation kernel is generated based on the dominant direction of the error corresponding to the aforementioned subspace. In order to ensure that the perturbation has spatial sensitivity and differences in the direction of action, this step performs eigenvalue decomposition on the covariance structure of each subspace, and extracts the principal component direction as the basis for constructing the perturbation basis vector. This direction vector not only carries the structural trend information within the subspace, but also implies the directional semantics of the subspace that is most likely to form a reconstruction vulnerability during the dimensionality reduction process, and is therefore very suitable as a perturbation direction expansion path.
[0112] In this embodiment, during the actual construction process, the perturbation kernel is not only determined by the main direction, but also a Gaussian perturbation term is added to enhance the instability and unpredictability of the perturbation response. The Gaussian perturbation term can be set to a group of random variables with zero mean and controlled variance to simulate nonlinear perturbation behavior while avoiding the reversible derivation of the perturbation trajectory. Through the superposition of this directional perturbation basis and the perturbation noise term, each perturbation matrix is not only separated from other matrices in terms of scope of action, but also has structural differences in the perturbation path, further strengthening the indecomposability of the embedded representation structure after the perturbation.
[0113] As an example, in S2.2.3.4, a set of independent tuning rules and tuning amounts are configured for each perturbation matrix based on the error statistical characteristics of each subspace and the structural characteristics of its corresponding perturbation kernel. The tuning rules include whether the perturbation direction can be adaptively adjusted, whether the perturbation frequency is activated based on feedback control, and whether the perturbation response introduces cross-matrix linkage feedback and other strategic mechanisms. The tuning amount involves quantifiable parameters such as the perturbation intensity boundary, the perturbation step coefficient, and the perturbation probability threshold. These configurations not only ensure the heterogeneity of each perturbation matrix when performing perturbation tasks, but also enable the entire perturbation system to have dynamic response capabilities and adaptability.
[0114] In this embodiment, to maintain the heterogeneity of the control mechanisms of each perturbation matrix during operation and maximize the inconsistency of the perturbation trajectory, the system configures a set of independent tuning rules and corresponding tuning variables for each characteristic subspace, specifically the error information density, error distribution type (whether it is concentrated or has boundary fluctuations), and the directional stability and perturbation sensitivity of its main perturbation kernel. This constitutes a dynamic adjustment mechanism for the perturbation matrix. This configuration mechanism neither uses template reuse nor fixed parameter input. Instead, it uses statistical data collected during the previous dimensionality reduction and reconstruction stages to perform driven parameter solution and selection. Its core goal is to ensure that the response behavior of each perturbation path is both diverse and does not introduce structural conflicts or spatial coupling.
[0115] Specifically, during the configuration of the tuning rules, the system first evaluates the response space of the main perturbation kernel's directional offset based on its directional offset stability. Specifically, if the angle between the main direction and the feedback vector oscillates frequently within a certain range during a continuous perturbation response, it indicates phase drift in the perturbation direction and requires a direction update mechanism. At this point, the tuning rules activate an adaptive perturbation direction adjustment strategy. When the feedback vector amplitude exceeds the upper bound of the projection intensity of the current main perturbation direction vector, the perturbation direction reconstruction function is triggered, continuously rotating the main perturbation kernel's direction around the axis of the feedback vector. The rotation amplitude increases proportionally with the angle value. On the other hand, if the subspace error response exhibits a periodic amplification-suppression trend and the perturbation response function converges slower than expected, the system identifies it as a low-sensitivity, high-inertia subspace and activates a perturbation frequency delay activation mechanism in the tuning rules. This mechanism triggers a new perturbation response after a fixed number of rounds or when the error increases significantly, ensuring that frequent perturbations do not degrade perturbation quality. Furthermore, in the subspace where multiple perturbation matrices have covariance structures related, if a subspace has been indirectly affected by non-local perturbations, the system will dynamically detect the variance mutation of its local perturbation response function and activate the cross-feedback linkage mechanism in the tuning rule, writing the second-order perturbation responses from other perturbation matrices into the current perturbation basis vector through the coupling function to perform a dynamic gain rotation, thereby realizing asynchronous interference compensation between different perturbation kernels.
[0116] In setting the tuning amount, the system first constructs a perturbation intensity benchmark interval based on the global intensity average and local fluctuation amplitude of the subspace error. The upper and lower boundaries of this interval correspond to the maximum error signal growth and minimum response compression amplitude of the current subspace after being perturbed, thereby limiting the dynamic domain of the perturbation amplitude. If the subspace still maintains a low amplitude response under multiple perturbations, the system automatically tightens its perturbation intensity upper limit and raises the lower limit to enhance its perturbation expression density. In terms of perturbation step size, the initial setting of the step size coefficient refers to the projection density of the main perturbation direction in the embedded space, that is, the cosine value of the angle between the perturbation basis vector and the current vector group. The closer the value is to 1, the more redundant the current perturbation is in the space, and the step size should be appropriately reduced to increase the perturbation frequency. Otherwise, the step size should be accelerated to quickly cross the directional damping zone. Similarly, the perturbation probability threshold is not a static constant but is dynamically updated based on the sensitivity of the main perturbation kernel's direction to the cross-feedback vector. Before a perturbation operation, the system will use a perturbation lookahead calculation (i.e., a small projection of the current perturbation direction to simulate the response trend) to determine whether to trigger the perturbation operation. If the fluctuation value of the pre-response function falls below the set perturbation threshold, the current perturbation behavior is suppressed. This mechanism ensures that perturbations are not frequently executed in inefficient areas, thereby saving computing resources and improving the expected response value of perturbation operations.
[0117] Furthermore, in order to prevent the disturbance matrix from falling into a stable mode during long-term operation, the tuning mechanism is also equipped with a periodic disturbance self-check module: when the disturbance increment of a disturbance matrix within multiple subspace periods converges to zero or near-zero, the system will consider that its disturbance kernel has entered a high stability zone or encountered feedback saturation, and will automatically reset its current tuning amount or re-randomize the disturbance direction, thereby breaking the inertial evolution trend of the disturbance path.
[0118] Notably, because the tuning rules and tuning amounts of each perturbation matrix are non-shared, there is no risk of unifying or collaborating during the overall perturbation process. Instead, through differentiated tuning strategies, the perturbations between the matrices exhibit asymmetric responses and unequal strength distributions, making it impossible to use any perturbation trajectory in the embedded representation to infer the overall field information, significantly enhancing the irreversibility of semantic modeling.
[0119] S2.2.4: Perturbing the corresponding feature subspace according to the perturbation matrix, and generating a local perturbation expression based on the perturbation response, wherein the perturbation is also cross-acted based on the perturbation responses of the feature subspaces corresponding to other perturbation matrices;
[0120] Specifically, the perturbation operation is not limited to performing random perturbations within the local feature subspace, but also introduces a cross-subspace cross-feedback mechanism. The goal is to make each perturbation behavior simultaneously affected by the perturbations in other subspaces, thereby forming a dynamic perturbation coupling structure and improving the unpredictability and irreversibility of the overall expression. Traditional piecewise perturbation methods often have the security risk of structural disassembly and perturbation response rollback due to the lack of cross-space feedback.
[0121] In this embodiment, the perturbation matrix first performs a perturbation operation on its corresponding feature subspace, including feature shifting along the perturbation kernel, injecting Gaussian noise with a specific covariance structure, and applying dynamic weighting of the perturbation intensity. The system then collects the differences, covariances, and relative angles between the current perturbation response vector and the perturbation response vectors of other subspaces to construct a cross-feedback vector, which is used to further guide the adjustment of the local perturbation direction.
[0122] Furthermore, each perturbation matrix performs updates such as directional rotation, perturbation gain scaling, or perturbation condition reset during the next perturbation cycle, based on the angle calculated between the cross-feedback vector and the current perturbation direction. The resulting perturbation response is no longer the result of a single perturbation direction, but rather a composite of responses in a coupled field of multiple perturbations, forming a set of local perturbation expressions that are nested, irreversible, and dynamically adaptable.
[0123] In an example, the specific steps of perturbing the feature subspace are as follows:
[0124] S2.2.4.1: Perform a perturbation operation on the feature embedding vectors in each feature subspace according to the perturbation matrix to generate a preliminary perturbation response, wherein the perturbation operation includes feature shifting and injecting random Gaussian noise;
[0125] Specifically, the perturbation kernel, structured in the previous stage, is applied to the actual subsegment of the embedding vector space. Because each feature subspace aggregates a similar set of principal dimensions based on dimensionality reduction and error partitioning, local perturbations are more likely to concentrate their impact on potentially high-semantic-strength regions. By performing biased shifts based on the perturbation kernel's direction and superimposing a random Gaussian perturbation term with adjustable covariance, the embedding vector can be deviated from its original stable direction, forming the perturbed initial representation.
[0126] In this embodiment, the perturbation operation adopts the following strategy: first, the perturbation direction vector is obtained according to the main perturbation kernel in the perturbation matrix, and then a directed amplitude offset is performed in this direction. The offset can be a fixed step size or dynamically scaled according to the local intensity function of the error in the subspace. At the same time, an additional set of Gaussian noise vectors with zero mean distribution and non-fixed dimensional co-correlation structure are introduced. These noise terms act on the current feature embedding through local injection to ensure that the final perturbation expression contains both directional perturbations and unpredictable perturbations. The effect is to form a noisy perturbation response, which contains irreversible perturbation trajectories.
[0127] Furthermore, the perturbation operation can be completed in one go, or multiple rounds of superimposed perturbations can be used, wherein the perturbation intensity of each round decreases, but the perturbation direction is continuously corrected with feedback to achieve stable coverage of multi-dimensional perturbations.
[0128] S2.2.4.2: For any perturbation matrix, when performing the perturbation operation, generate a cross-feedback vector based on the preliminary perturbation response of the corresponding characteristic subspace of the other perturbation matrices;
[0129] Specifically, if all perturbation matrices operate independently, separable perturbation intervals may form between subspaces, allowing for potential deconstruction using reverse engineering techniques. By introducing cross-feedback, the closed nature of the perturbation path can be effectively broken, allowing the perturbation direction to be controlled by multi-space interference inputs, thereby improving the unpredictability of the perturbation.
[0130] In this embodiment, the cross-feedback vector is generated as follows: After each perturbation matrix generates its own preliminary perturbation response, it performs a differential calculation between the post-perturbation vector and the pre-perturbation vector to obtain a perturbation gradient response vector. Simultaneously, all other perturbation matrices also generate corresponding gradient responses. For any perturbation matrix, the perturbation direction update in the current cycle will take into account these gradient signals from other subspaces, constructing a fused perturbation direction vector that includes weight adjustment factors.
[0131] Furthermore, to prevent all perturbation vectors from converging due to feedback merging, a covariance weighting coefficient or cosine similarity adjustment coefficient is used during the fusion process between feedback vectors to ensure selective directional updates. The essence of this feedback mechanism is perturbation driving perturbation, with perturbations in different subspaces influencing each other, forming a perturbation chain.
[0132] In practice, to construct this perturbation chain, each perturbation matrix must be configured with a perturbation feedback monitoring period and a cross-signal fusion strategy. The monitoring period can be fixed or controlled by dynamic entropy, and the feedback fusion strategy can be angle-driven, intensity-driven, or principal component-driven.
[0133] S2.2.4.3: Adjusting the perturbation direction of the current perturbation matrix according to the cross-feedback vector to generate a secondary perturbation response, specifically comprising: calculating the angle between the cross-feedback vector and the characteristic subspace direction vector;
[0134] Rotating the disturbance direction of the current disturbance matrix around the cross feedback vector by the angle to generate an updated disturbance direction;
[0135] Specifically, since the perturbation direction may be inconsistent with the semantic direction of the feature subspace, direct replacement will lead to perturbation imbalance or failure to maintain subspace coupling, so it is more reliable to use a vector rotation mechanism.
[0136] In this embodiment, the perturbation direction is rotated as follows: First, the angle between the current perturbation direction vector and the cross-feedback vector is calculated, and the angle value is obtained through the vector dot product method. Then, with the cross-feedback vector as the rotation axis, the current perturbation direction is rotated around this axis by a specified angle using a three-dimensional vector rotation model (Rodrigues formula or quaternion rotation). The rotation angle can be full angle or scaled to achieve perturbation deflection. The rotated perturbation direction is used for the next perturbation operation, completing one perturbation direction iteration with feedback.
[0137] Furthermore, an adaptive tuning mechanism can be introduced to adjust the rotation angle, where the rotation amplitude is determined based on the variance of the cross-feedback or its orthogonality with the current perturbation direction. If the feedback vector recurs in multiple subspaces, angle suppression or perturbation redirection can be performed to avoid perturbation resonance.
[0138] S2.2.5: Nest and combine the local perturbation expressions to generate a semantic embedding representation;
[0139] Specifically, the local perturbation representations of each feature subspace after perturbation alone cannot constitute a complete semantic model; they require structural fusion and semantic integration. Simple splicing cannot express the correlations between perturbations across spaces, so a nested combination mechanism is needed to construct a semantic embedding representation that is hierarchical, combinatorial, and contextually coupled.
[0140] In this embodiment, a hierarchical nested combination strategy is employed, dividing each local perturbation representation into multiple nested levels based on perturbation intensity, structural complexity, and subspace semantic weight. The outer layer primarily carries low-risk semantic response areas, while the core layer aggregates structural information for high-intensity perturbation areas. This nested structure enables subsequent response modules to perform structural unpacking or response regulation based on access requirements, and also provides nested path support for access entropy determination.
[0141] S2.3: Encapsulate the semantic embedding representation into a semantic representation model;
[0142] Specifically, although the semantic embedding representation after the disturbance processing is irreversible, it has not yet formed a complete responsive model structure. In order to make the semantic embedding representation have the ability to respond to subsequent access, this embodiment performs structured encapsulation on it to generate a unified semantic representation model. The encapsulation process includes semantic structure layer registration, index path binding and access control identifier injection. In the encapsulation structure, each semantic unit or semantic fragment will be mapped to a set of response dimension labels for semantic intent matching in the access request. At the same time, in order to avoid the leakage of the dimensional space mapping path within the structure, this embodiment does not retain the explicit field to semantic dimension bidirectional mapping table, but only retains the one-way index path from the response path to the original field, and the index ontology is independent of the main structure of the model.
[0143] As an example, an embodiment of the present invention provides a method for managing business candidate information. When an external access request is received, the method further includes:
[0144] S4: parsing the external access request;
[0145] Specifically, the semantic representation model of business candidates cannot be accessed through traditional field matching methods because it no longer contains traceable field information in its structure. In order to achieve semantic-level response generation, it is necessary to first parse its semantic intent after the access request arrives, so as to perform structural-level semantic matching in the representation model. In this embodiment, the external access request usually includes information such as the requester identifier, the requested task type, the target position feature description, or the natural language query intent. The parsing operation is implemented by constructing an access intention vector, which is generated using the same encoding specifications as the semantic representation model to ensure that the two are structurally comparable. The construction of the access intention vector may include: context segmentation of the natural language description, label extraction and structure mapping, and finally generating a set of nested semantic identification vectors to be used as input for the subsequent matching process.
[0146] S5: matching semantic response dimensions from the semantic representation model according to the parsing result;
[0147] Specifically, since the semantic representation model is constructed using an irreversible perturbation strategy, the field information has been disturbed, compressed, and nested, leaving only the responsive semantic hierarchical structure. Therefore, the matching process needs to be calculated layer by layer based on the semantic similarity between the access intention vector and the semantic nested structure. During implementation, the access intention vector is calculated for similarity with all nested nodes in the candidate's semantic representation model. This process uses indicators such as vector space cosine distance, nested path overlap rate, and semantic center offset to construct a multi-factor matching scoring function. When multiple semantic nodes meet the similarity threshold, the candidate set can be streamlined according to the path hierarchy priority or structural coupling degree, and finally a set of response dimension node sets are output. The response dimension set represents the semantic areas that can be activated in the semantic model for this access request. These areas will serve as dimension indexes in the response generation phase to trigger subsequent field information data query paths.
[0148] S6: Generate a response result according to the semantic response dimension, wherein the response result is queried in the field information data according to the index path;
[0149] Specifically, since the semantic model and the original field information adopt a separate storage mechanism, the model body does not contain field information data. Only after the semantic response is triggered, the corresponding fragment in the field cache is queried through the structural index path. In this embodiment, each semantic nested node will be bound to a set of field index identifiers during construction. The identifier will not be directly disclosed, but can be used to drive the field query module when the semantic dimension is activated. During the response process, the field information in the restricted cache is accessed according to the field index path corresponding to each node in the semantic response dimension set, and the minimum coverage fragment is extracted for structured output. The query process can set a field granularity control policy to limit security policies such as the maximum number of response fields and the maximum number of field characters to prevent inference of field content through multiple rounds of calls.
[0150] S7: After the response is completed, the semantic representation model is encapsulated according to the access control configuration;
[0151] Specifically, in order to prevent the access results from being iteratively extracted from the original model structure by multiple rounds of calls, the semantic representation model needs to be dynamically encapsulated. In this embodiment, whether to trigger the encapsulation strategy and select the encapsulation method are determined based on the coverage of the semantic response, the access level of the access subject and the access entropy value of this request. The encapsulation strategy may include: structural compression encapsulation, that is, deleting unactivated semantic branches and reconstructing the response hierarchy tree; summary encapsulation, that is, converting the response path structure into a semantic label index description to hide the structure depth; perturbation enhancement encapsulation, that is, applying a slight perturbation to the responded semantic node to prevent the next round of requests from repeatedly accessing the same response path. The calculation of the access entropy value is based on a comprehensive evaluation of parameters such as the response path length, semantic distribution density, and the number of access history, and is used to dynamically determine the risk level after the response. The encapsulated semantic model will replace the original model for subsequent semantic interactions to ensure the closed-loop use and dynamic self-protection of the model structure.
[0152] The specific steps of S5 are as follows:
[0153] S5.1: Construct access intention vector;
[0154] Specifically, in the technical architecture of this invention, because the candidate semantic representation model has already disrupted the direct mapping relationship between fields and semantic expressions through a disruption strategy, visitors cannot directly retrieve data by field name or field value. Therefore, the query target content in the external access request must be converted into a semantic expression vector that can be compared with the semantic nested structure. This semantic expression vector is the access intention vector, whose task is to extract semantic key information from natural language or structured input and construct a comparable semantic representation using an encoding strategy that is the same as or compatible with the nested structure in the semantic model.
[0155] In this embodiment, the construction process of the access intent vector includes the following technical steps: first, the original request text is parsed, and a pre-trained semantic recognition model is used to perform part-of-speech tagging, named entity recognition, and context window extraction to extract semantically targeted keyword phrases or labels; then, the extracted information is embedded in the vector space, and the encoding method used is consistent with the field embedding structure in the candidate semantic model, such as using a Transformer encoder fine-tuned in the domain or a dedicated semantic nested encoder. The constructed access intent vector not only retains the semantic core of the original request, but also retains the structural relationship between words and context information, thereby providing an accurate semantic mapping basis for subsequent semantic matching.
[0156] Furthermore, during the access intention vector construction phase, a structural entropy adjustment factor can be introduced based on the requester's identity, access frequency, and organizational role to adjust the entropy weight of the final generated semantic vector, thereby enhancing the recognition ability of the response area in the semantic overlapping area, while suppressing the repeated activation of frequently called semantic paths in the model and improving the balance of call distribution.
[0157] S5.2: Matching the access intention vector with the semantic embedding representation in the semantic representation model, wherein the matching includes semantic similarity calculation and structural path intersection;
[0158] Specifically, the candidate semantic representation model is not generated by arranging the original fields in a fixed position. Instead, it generates a nested structure of cross-representations in multiple subspaces through a perturbation matrix. Each semantic node represents a combination of several fields and their semantic fusion after undergoing irreversible transformations. Therefore, the matching process must be based on the similarity between vector representations in the semantic space and the path relationships formed by the nested structure in the model to accurately locate the semantic region corresponding to the access intent.
[0159] In this embodiment, the semantic matching operation first performs a semantic similarity calculation, that is, comparing the access intent vector with all nested nodes in the candidate person's semantic model one by one. The similarity evaluation method adopts a multi-factor calculation strategy, which not only includes the basic cosine similarity or Euclidean distance measurement, but also introduces structural context alignment evaluation, such as the hierarchical position of the node in the nested structure, the length of the parent-child relationship graph path, etc., to form a composite matching scoring mechanism. In order to improve the fault tolerance and sensitivity of the matching, a fuzzy semantic expansion strategy is also introduced. When the nested distribution of some semantic words does not completely overlap, the response layer is supplemented by guided weight shifting.
[0160] After completing the semantic similarity scoring, the set of candidate nodes with matching scores within the threshold is then used to determine structural path intersection. Path intersection refers to whether the inferred path mapped by the access intent vector overlaps with the semantic paths between nested nodes in the model. This includes factors such as path prefix consistency, path nesting level alignment, and path topology matching. In cases where semantic similarity is low but the structural paths are very close, the matching enhancement strategy is triggered, treating them as valid intersection nodes to improve matching coverage.
[0161] S5.3: Filter the semantic nodes whose semantic similarity is greater than a preset threshold or whose path overlap is greater than a preset condition from the matching results;
[0162] Specifically, to ensure the technical rigor of the selection of semantic response dimensions and avoid over-activation, the system must perform a refined screening of all preliminary matching nodes to output a set of executable semantic nodes. This set constitutes the response target area for this access request, and its screening criteria should comprehensively consider two core indicators: semantic similarity and structural overlap.
[0163] In this embodiment, the screening process implements the following strategy: the system pre-sets a similarity threshold and a path overlap threshold, sorts the similarity scores of each preliminary matching node, and directly includes nodes with scores above the similarity threshold in the response set. For nodes with similarity scores below the similarity threshold, if their path overlap index exceeds the path overlap threshold (for example, if the path prefix length ratio exceeds a certain ratio or the structural hierarchy is completely aligned), they are also determined to have a structurally dependent semantic match and are included in the response set.
[0164] Furthermore, the system can perform a semantic sparsification operation on the matching result set to prevent multiple nodes with similarity but overlapping fields from being activated simultaneously, which would lead to redundant response data. Sparsification is based on a cross-analysis of the fields covered by the semantic nodes, retaining the nodes with the highest semantic information density and removing nodes with marginal overlap.
[0165] S5.4: Using the index dimension set covered by the semantic node as the semantic response dimension;
[0166] Specifically, for each nested semantic node in the candidate semantic representation model, a corresponding index path mapping is established during construction. This mapping points to some or all field fragments within the field information data. In the model itself, these paths do not directly expose the field content, but serve as an intermediate bridge between the response logic and the field information data during the semantic response process. Therefore, using the set of index paths covered by the selected semantic nodes as the semantic response dimension is a key mechanism for achieving the compatibility of "structural response + data isolation."
[0167] In this embodiment, the system extracts the field index path identifiers recorded during the model building phase for each node from the selected semantic node set, and forms an index dimension set. This set guides subsequent field information data access queries and also serves as a structural description of the semantic response results, helping the caller understand the source of the response semantics.
[0168] To ensure minimal data access availability, the system implements a minimum coverage optimization strategy for index dimension sets, activating only the minimum subset of fields necessary for the access, thus avoiding invalid responses. This set can be labeled with a lifecycle throughout the call chain, limiting its validity to the current access task. This prevents semantic dimension information from being retained for extended periods, potentially leaking model structure.
[0169] The access control configuration includes an access entropy adjustment policy, and the access entropy adjustment policy includes:
[0170] S7.1: Calculate an access entropy value, where the access entropy value is calculated based on the nesting level depth of the semantic response dimension hit by the access request in the semantic representation model, the number of response dimensions, the structural path density, and the coverage ratio of the original field index path;
[0171] S7.2: Select an encapsulation strategy according to the access entropy value, wherein the encapsulation strategy includes structure compression encapsulation, summary encapsulation and scrambled encapsulation.
[0172] A business candidate information management and storage system, the system comprising:
[0173] The modeling module is used to obtain field information data of business candidates and build an irreversible vectorized representation model containing a semantic nested structure;
[0174] The perturbation module is used to perform random dimensionality reduction, error analysis, and cross-control of multiple perturbation matrices on the initial representation vector to generate an irreversible semantic embedding representation;
[0175] A storage module is used to store the semantic representation model and field information data separately and establish an index path connection structure;
[0176] The access module is used to parse external access requests, match semantic response dimensions, and generate response results through index paths;
[0177] The encapsulation module is used to calculate the access entropy value according to the access entropy adjustment strategy, and select the corresponding encapsulation method to perform structural compression, summary and scrambling operations on the semantic representation model.
[0178] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A method for managing and storing business candidate information, characterized in that: The information storage method includes: Acquiring multiple field information data of the business candidate, performing feature extraction and encoding on the multiple field information data, and constructing an initial representation vector group; Performing irreversible semantic modeling on the initial representation vector group according to a preset perturbation strategy to generate a semantic representation model, wherein the irreversible semantic modeling includes setting the initial representation vector group and reconstructing the error distribution, obtaining a feature subspace according to a main dimension cluster of the error, and initializing a perturbation matrix for each feature subspace, wherein the tuning rules and tuning amounts of the perturbation matrices are different from each other, wherein the perturbation matrix acts on the corresponding feature subspace and generates a cross-feedback vector according to the perturbation response of the feature subspace, generating a secondary perturbation response according to the cross-feedback vector to update the corresponding feature subspace, generate a semantic embedding representation, and generate a semantic representation model according to the semantic embedding representation; The semantic representation model is stored in a preset information database.
2. A method for managing and storing business candidate information according to claim 1, characterized in that: The disruption strategies include: Perform the first round of random dimensionality reduction on the initial representation vector group to generate the first round of dimensionality reduction embedding vector; Performing an inverse reconstruction calculation on the first round of dimensionality reduction embedding vectors to obtain a reconstruction error distribution feature between the vectors and the initial representation vector group; Initializing at least two perturbation matrices according to the reconstruction error distribution characteristics, respectively acting on different feature subspaces of the embedding vector obtained by the first round of dimensionality reduction; Perturbing the corresponding feature subspace according to the perturbation matrix, and generating a local perturbation expression according to the perturbation response, wherein the perturbation is also cross-acted based on the perturbation response of the feature subspace corresponding to other perturbation matrices; The local perturbation expressions are nested and combined to generate a semantic embedding representation.
3. A method for managing and storing business candidate information according to claim 2, characterized in that: Initializing at least two disturbance matrices according to the reconstruction error distribution characteristics, including: Calculating the error intensity distribution of the reconstruction error distribution feature in different feature dimensions; Dividing the first round of dimensionality reduction embedding vector into at least two feature subspaces according to the error intensity distribution, wherein each feature subspace corresponds to a main dimension cluster of error intensity; Initializing a main perturbation kernel of a perturbation matrix according to the direction vector of the main dimension cluster, wherein the main perturbation kernel includes a perturbation basis vector constructed along the direction vector and a Gaussian perturbation term; According to the statistical characteristics of the error in the characteristic subspace and the structural characteristics of the main perturbation kernel, the tuning rule and tuning amount of the perturbation matrix are determined.
4. A method for managing and storing business candidate information according to claim 3, characterized in that: The tuning rules include a disturbance direction update method, an amplitude gain strategy, and a disturbance activation condition, and the tuning variables include a disturbance intensity range, an update step, and a disturbance probability threshold.
5. A method for managing and storing business candidate information according to claim 2, characterized in that: Perturbing the corresponding feature subspace according to the perturbation matrix includes: Performing a perturbation operation on the feature embedding vectors in the respective feature subspaces according to the perturbation matrix to generate a preliminary perturbation response, wherein the perturbation operation includes feature shifting and injecting random Gaussian noise; For any perturbation matrix, when performing the perturbation operation, a cross-feedback vector is generated according to the preliminary perturbation response of the characteristic subspace corresponding to the other perturbation matrix; The disturbance direction of the current disturbance matrix is adjusted according to the cross-feedback vector to generate a secondary disturbance response.
6. A method for managing and storing business candidate information according to claim 5, characterized in that: Adjusting the disturbance direction of the current disturbance matrix according to the cross feedback vector includes: Calculating the angle between the cross feedback vector and the feature subspace direction vector; The disturbance direction of the current disturbance matrix is rotated around the cross feedback vector by the angle to generate an updated disturbance direction.
7. A method for managing and storing business candidate information according to claim 1, characterized in that: When an external access request is received, the information storage method further includes: Parsing the external access request; matching semantic response dimensions from the semantic representation model according to the parsing results; Generate a response result according to the semantic response dimension, wherein the response result queries the field information data according to the index path; After the response is completed, the semantic representation model is encapsulated according to the access control configuration.
8. A method for managing and storing business candidate information according to claim 7, characterized in that: Matching the semantic response dimension from the semantic representation model according to the parsing result includes: Build access intent vector; Matching the access intention vector with the semantic embedding representation in the semantic representation model, wherein the matching includes semantic similarity calculation and structural path intersection; Filter semantic nodes whose semantic similarity is greater than a preset threshold or whose path overlap is greater than a preset condition from the matching results; The index dimension set covered by the semantic node is used as the semantic response dimension.
9. A method for managing and storing business candidate information according to claim 7, characterized in that: The access control configuration includes an access entropy adjustment policy, and the access entropy adjustment policy includes: Calculating an access entropy value, wherein the access entropy value is calculated based on the nesting level depth, the number of response dimensions, the structural path density, and the coverage ratio of the original field index path of the semantic response dimension hit by the access request in the semantic representation model; An encapsulation strategy is selected according to the access entropy value, wherein the encapsulation strategy includes structure compression encapsulation, summary encapsulation and scrambled encapsulation.
10. A business candidate personnel information management and storage system, used to implement a business candidate personnel information management and storage method according to any one of claims 1 to 9, characterized in that: The system comprises: The modeling module is used to obtain field information data of business candidates and build an irreversible vectorized representation model containing a semantic nested structure; The perturbation module is used to perform random dimensionality reduction, error analysis, and cross-control of multiple perturbation matrices on the initial representation vector to generate an irreversible semantic embedding representation; A storage module is used to store the semantic representation model and field information data separately and establish an index path connection structure; The access module is used to parse external access requests, match semantic response dimensions, and generate response results through index paths; The encapsulation module is used to calculate the access entropy value according to the access entropy adjustment strategy, and select the corresponding encapsulation method to perform structural compression, summary and scrambling operations on the semantic representation model.
Citation Information
Patent Citations
Text recognition method based on language model
CN120012079A
Private data protection method and device, storage medium and computer program product
CN119323051A
Human-post matching recommendation method based on BERT and latent semantic algorithm model
CN119377490A