Business candidate person and member information management and storage method and system
By building an irreversible semantic representation model and introducing a multi-perturbation matrix, the problem of sensitive information easily leaked in the candidate information management system is solved, and access controllability and security are achieved.
Patent Information
- Application Number
- CN202510885151.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-30
AI Technical Summary
In the prior art, the candidate information management system lacks structural irreversibility, disturbance anti-reduction ability and access controllability while its semantic response capabilities, resulting in susceptible leakage of sensitive information and privacy violations.
By constructing the initial representation vector group, irreversible semantic modeling is used to use a scramble strategy to generate a semantic embedded representation, and separate it from the original field data, it is accessed only through the index path, and a multi-perturbation matrix is introduced to act on different feature subspaces, combining the tuning rule differences and the cross-feedback mechanism to interrupt the reversible field path.
It improves the control security of sensitive information and the controllability of model calls, ensuring that candidate field information is not restored during access, and prevents data leakage.
Smart Images

Figure CN120372669A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage, and particularly to a method and system for managing and storing information of business candidate personnel. Background Art
[0002] In existing talent information management systems, the field information of candidate personnel is usually stored centrally in a structured form and retrieved and analyzed through field-level indexes. However, such field information often contains sensitive content such as names, contact information, work resumes, ability tags, etc. Once illegally accessed or the desensitization fails, it is extremely easy to cause data leakage, privacy infringement, and organizational security risks. To improve the access efficiency and intelligence level, some systems introduce a vector representation mechanism, but most solutions still retain the field-level semantic boundaries or have the ability to restore fields, and cannot block the mapping link from semantic expression to field content from the technical path.
[0003] For example, the Chinese patent application with the publication number CN120012079A provides a text recognition method based on a language model, including obtaining a text data stream to be detected in a target application program, including user input, interface display, and network transmission text; preprocessing the text data stream to generate a multi-dimensional semantic vector sequence; modulating the features of the semantic vector sequence through a dynamic weight allocation module to generate a modulated feature matrix; inputting a logic judgment engine to perform multi-level conditional combination analysis to generate a risk probability distribution; demodulating the risk probability distribution to generate an interpretable risk feature vector; performing cross-validation based on the risk feature vector and a preset threshold set to generate a final risk determination result; when it is determined as a high-risk, triggering a multi-layer protection mechanism, including content replacement, session interruption, and security alarm. The invention can improve the accuracy and interpretability of the system and ensure the security and user experience of the target application program.
[0004] All of the above prior arts have the problems proposed in this background art: there is a lack of a method for managing and storing information of candidate personnel that has structural irreversibility, anti-restoration ability to perturbations, and access controllability while ensuring semantic response ability. To solve the above problems, this application designs a method and system for managing and storing information of business candidate personnel. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method and system for managing and storing information of business candidate personnel in view of the deficiencies of the prior art, obtaining multiple field information data of business candidate personnel, and constructing an initial representation vector group; performing irreversible semantic modeling according to a predefined perturbation strategy to generate a semantic embedding representation and encapsulating it into a semantic representation model, where the perturbation strategy includes dividing the feature subspace, setting perturbation matrices with different tuning rules, and asymmetric coupling perturbation based on cross-feedback; the semantic model is stored separately from the original field data, and the field is only controlled and accessed through an index path to achieve data response rather than exposure. This method improves the control security of sensitive information and the controllability of model invocation.
[0006] To achieve the above object, the present invention provides the following technical solutions: A method for managing and storing information of business candidate personnel, the information management and storage method comprising: Obtaining multiple field information data of the business candidate personnel, performing feature extraction and encoding on the multiple field information data, and constructing an initial representation vector group; Performing irreversible semantic modeling on the initial representation vector group according to a predefined perturbation strategy, where the perturbation strategy includes: Setting the initial representation vector group and performing a reconstruction error distribution, and obtaining a feature subspace according to the main dimension cluster of the error; Initializing a perturbation matrix for each feature subspace, and the tuning rules and tuning amounts of each perturbation matrix are different from each other; The perturbation matrix acts on the corresponding feature subspace and generates a cross-feedback vector according to the perturbation response of the feature subspace, generating a secondary perturbation response according to the cross-feedback vector to update the corresponding feature subspace, generating a semantic embedding representation, and generating a semantic representation model according to the semantic embedding representation; Storing the semantic representation model in a preset information database.
[0007] The perturbation strategy includes: Performing a first-round random dimensionality reduction operation on the initial representation vector group to generate a first-round dimensionality reduction embedding vector; Performing reverse reconstruction calculation on the first-round dimensionality reduction embedding vector to obtain a reconstruction error distribution feature between the first-round dimensionality reduction embedding vector and the initial representation vector group; Initializing at least two perturbation matrices according to the reconstruction error distribution feature, and respectively acting on different feature subspaces of the first-round dimensionality reduction embedding vector; Perturbing the corresponding feature subspace according to the perturbation matrix, and generating a local perturbation expression according to the perturbation response, where the perturbation also cross-acts based on the perturbation responses of other perturbation matrix corresponding feature subspaces; Nesting and combining the local perturbation expressions to generate a semantic embedding representation.
[0008] Initialize at least two perturbation matrices according to the reconstructed error distribution characteristics, including: Calculate the error intensity distribution of the reconstructed error distribution characteristics on different feature dimensions; Divide the first-round dimensionality-reduced embedding vectors into at least two feature subspaces according to the error intensity distribution, where each feature subspace corresponds to a main dimension cluster of an error intensity; Initialize the main perturbation kernel of the perturbation matrix according to the direction vector of the main dimension cluster, where the main perturbation kernel includes perturbation basis vectors constructed along the direction vector and Gaussian perturbation terms; Determine the tuning rules and tuning amounts of the perturbation matrix according to the statistical characteristics of the errors in the feature subspace and in combination with the structural characteristics of the main perturbation kernel.
[0009] The tuning rules include a perturbation direction update method, an amplitude gain strategy, and a perturbation activation condition, and the tuning amounts include a perturbation intensity range, an update step size, and a perturbation probability threshold.
[0010] Perturb the corresponding feature subspace according to the perturbation matrix, including: Perform a perturbation operation on the feature embedding vectors in their respective feature subspaces according to the perturbation matrix to generate a preliminary perturbation response, where the perturbation operation includes feature offset and injection of random Gaussian noise; For any perturbation matrix, when performing the perturbation operation, generate a cross-feedback vector according to the preliminary perturbation responses of the feature subspaces corresponding to other perturbation matrices; Adjust the perturbation direction of the current perturbation matrix according to the cross-feedback vector to generate a secondary perturbation response.
[0011] Adjust the perturbation direction of the current perturbation matrix according to the cross-feedback vector, including: Calculate the angle between the cross-feedback vector and the direction vector of the feature subspace; Rotate the perturbation direction of the current perturbation matrix around the cross-feedback vector by the angle to generate an updated perturbation direction.
[0012] When an external access request is received, the information storage method further includes: Parse the external access request; Match a semantic response dimension from the semantic representation model according to the parsing result; Generate a response result according to the semantic response dimension, where the response result is queried from the original field data according to an index path; After the response is completed, encapsulate the semantic representation model according to the access control configuration.
[0013] Matching semantic response dimensions from the semantic representation model according to the parsing result includes: Construct an access intention vector; Match the access intention vector with the semantic embedding representation in the semantic representation model, where the matching includes semantic similarity calculation and structural path intersection; Filter semantic nodes with semantic similarity greater than a preset threshold or path coincidence degree greater than a preset condition from the matching results; Use the set of index dimensions covered by the semantic nodes as the semantic response dimensions.
[0014] The access control configuration includes an access entropy adjustment policy, and the access entropy adjustment policy includes: Calculate the access entropy value, where the access entropy value is calculated based on the nested level depth, the number of response dimensions, the structural path density, and the coverage ratio of the original field index path in the semantic response dimensions hit by the access request; Select an encapsulation policy according to the access entropy value, where the encapsulation policy includes structural compression encapsulation, summarization encapsulation, and scrambling encapsulation.
[0015] A business candidate information management and storage system, the system includes: A modeling module, configured to obtain the field information data of business candidates and construct an irreversible quantization representation model including a semantic nested structure; A perturbation module, configured to perform random dimensionality reduction, error analysis, and multi-perturbation matrix cross-control on the initial representation vector to generate an irreversible semantic embedding representation; A storage module, configured to separately store the semantic representation model and the original field data and establish an index path connection structure; An access module, configured to parse an external access request, match semantic response dimensions, and generate a response result through the index path; An encapsulation module, configured to calculate the access entropy value according to the access entropy adjustment policy and select a corresponding encapsulation method to perform structural compression, summarization, and scrambling operations on the semantic representation model.
[0016] Compared with the prior art, the beneficial effects of the present invention are: The present invention converts the candidate field information into an irreversible semantic representation model, introduces a multi-perturbation matrix to act on different feature subspaces, combines the tuning rule differences and the cross-feedback mechanism, realizes the deep disruption of the mapping structure between the field information and the semantic expression, and effectively interrupts the reversible path of the field. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present invention will become more apparent: Figure 1 This is a schematic diagram of an exemplary application scenario for an embodiment of the present invention; Figure 2 This is a schematic flow diagram of a method for managing and storing information of business candidate personnel in an embodiment of the present invention; Figure 3 This is a schematic flow diagram of a data conversion method in an embodiment of the present invention. Detailed implementation manners
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0019] The "embodiment" mentioned in this article means that the specific features, structures, or characteristics described in conjunction with the embodiment may be included in at least one embodiment of the present application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art can explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0020] Please refer to Figure 1 , which is a schematic diagram of an exemplary application scenario provided by an embodiment of the present application.
[0021] As Figure 1 shown, the application scenario involves the sharing of talent models among multiple subsidiaries within a group, including Subsidiary A, Subsidiary B, and Subsidiary C, all of which have local candidate field data. To avoid direct upload of sensitive fields, each subsidiary converts the candidate field data into a semantic representation model through an irreversible modeling method.
[0022] Figure 1 It shows that Subsidiary A generates semantic representation model A, Subsidiary B generates semantic representation model B, and Subsidiary C generates semantic representation model C.
[0023] In one example, the above semantic representation models are all irreducible vector expression structures.
[0024] Figure 1 It shows that after each subsidiary completes semantic modeling locally, it uploads the generated semantic representation model to the talent reserve library of the headquarters platform. The headquarters platform centrally stores each semantic representation model for responding to access requests from within the platform or authorized systems.
[0025] Figure 1It shows that when the headquarters platform receives an access request, it matches the corresponding semantic representation model according to the content of the access request, and generates semantic response content based on the matching result, forming a response result to be returned to the access party. During the whole process, the candidate field data is not transmitted, stored or displayed on the headquarters platform, and the access request is only responded through an irreversible semantic representation model, ensuring the isolation and security of the field information.
[0026] In one example, the semantic representation model can be used by the headquarters platform to perform talent ability modeling and job matching across subsidiaries. Through the unified semantic nesting structure definition, the semantic models uploaded by different subsidiaries are comparable at the structural level, and the platform can calculate the semantic similarity between candidates and jobs based on the semantic dimension, which is used for business requirements such as job recommendation, echelon construction, and organizational ability prediction.
[0027] In one example, the platform can generate a candidate ability graph based on the aggregated semantic representation model. This graph does not rely on field-level attributes, but only on the semantic structure and feature transformation results, and is applicable to talent distribution analysis, ability gap identification, and organizational strategy assessment, with the adaptability across systems, languages, and organizations.
[0028] On the other hand, the irreversible modeling method can be deployed in a private isolated computing environment, only allowing local use of field data to complete the generation of semantic vectors, and immediately destroying the intermediate field cache after the model is generated, without generating the link risk of field-level data leakage, meeting the data compliance and cross-organizational model collaboration security requirements in special scenarios.
[0029] Next, in combination with the accompanying drawings, a method for managing and storing business candidate information provided by an embodiment of the present application will be introduced.
[0030] Please refer to Figure 2 , which is a schematic flow chart of a method for managing and storing business candidate information provided by an embodiment of the present application. Figure 2 The method shown can be applied to the group internal, government-enterprise platform or multi-source talent collaboration system to perform semantic modeling and access protection for candidate field data. Figure 2 The method shown includes the following S1 - S3, and the specific steps are as follows: S1: Obtain multiple field information data of business candidates; In this step, the field information data may include data content containing field information extracted from the enterprise internal database, human resources system or third-party candidate interface, and the fields include but are not limited to structured or semi-structured information such as name, education background, job experience, ability labels, etc.
[0031] S2: Convert the field information data into a semantic representation model; In this step, based on the predefined semantic modeling mechanism, feature extraction and encoding are performed on the field information data to construct an initial representation vector group; Subsequently, an irreversible modeling strategy is executed. This strategy includes random dimensionality reduction, reconstruction error calculation, perturbation matrix initialization, and cross-feedback perturbation to generate an irreversible quantization expression body with a semantic nested structure, where the irreversible quantization expression body is the semantic embedding representation; The semantic embedding representation is encapsulated into a semantic representation model. This model does not contain directly recognizable field semantic boundaries and is used to drive semantic layer data interaction when responding to access requests later.
[0032] S3: Store the semantic representation model separately from the field data and establish an access index path; In this step, the semantic representation model is bound with access control configuration and stored separately from the original field data. The field data is placed in a temporary access buffer and is only accessed restrictively through the index path according to the semantic response dimension when the access control conditions are met, ensuring that the candidate's field information is not leaked or reconstructed during the entire modeling and invocation process.
[0033] Although in the prior art, semantic modeling and feature vector expression have been widely applied to tasks such as talent recommendation, resume screening, and candidate ability matching, the current common practice still mainly relies on reducible structures or weak semantic isolation mechanisms. Even after desensitization, the vector structure itself can still be used to restore some or all of the original field information through means such as model reverse inference or field distribution backtracking. The typical feature of such technical solutions is that while their vector expressions retain field semantic boundaries, field context structures, and encoding paths, they fail to effectively break the corresponding relationship between the semantic structure and the original fields. Therefore, during the model invocation process, attackers can gradually restore the candidate's sensitive fields by frequently accessing the semantic interface to construct a dimension frequency map, a semantic similarity mapping, or a gradient reverse inference path, resulting in privacy leakage or resource competition risks.
[0034] In this embodiment, by constructing a random initial dimensionality reduction space and introducing a multi-perturbation matrix mechanism driven by error feedback, the system performs multiple rounds of asymmetric and structurally decoupled dimensionality reduction operations on the original field vector group, so that the finally generated semantic embedding model no longer retains the semantic boundaries between fields in terms of structure, and the vector representation itself no longer has a reversible mapping path. Compared with traditional embedding methods, the solution of this application breaks the stable mapping relationship between fields and semantic representations while maintaining the responsiveness of the model. It should be noted that the method of this application uses a cross-perturbation strategy between multiple matrices, so that each perturbation response is not only independent of the field source, but also forms a coupling with the subspace feedback of other perturbation matrices, thereby constructing an expression structure that is only responsive in local dimensions but not globally inferable. This structure inherently has the ability to resist reconstruction. Even if an attacker has some field labels and the semantic response results of the model, it is difficult to reconstruct the original field vectors or identify the sensitive entities corresponding to the semantic expression units.
[0035] Exemplarily, in a scenario of talent sharing within a group, multiple subsidiaries each have independent original field information of candidates. The traditional solution may upload the field desensitized and the generated embedding vectors to the headquarters platform. However, since these vectors still retain the field semantic path in terms of structure, there is a risk of being restored. In the implementation manner of this application, each subsidiary performs multi-matrix perturbation modeling on the field information locally and only uploads the irreversible semantic model to the headquarters. When the headquarters platform makes a matching request for a certain position, only the semantic response result can be obtained, and the response path is controlled by the access entropy, ensuring that even the headquarters platform cannot restore the candidate fields. In this way, the platform can complete semantic layer operations such as talent recommendation and ability analysis without involving any field leakage problems.
[0036] In this embodiment, since there are usually different semantic granularities, numerical types, and statistical distribution characteristics among field information data, directly performing vectorization processing may lead to insufficient semantic coupling or unbalanced embedding space dimensions, thereby affecting the quality of subsequent irreversible modeling. Therefore, before semantic modeling, this embodiment preferably performs structured preprocessing on the original field information data, specifically including field screening, type normalization, and multi-scale semantic reconstruction. Among them, field screening is used to eliminate field information with weak correlation or poor stability with the modeling task, such as extremely sparse labels, low-correlation evaluation items, etc., to reduce the interference caused by redundant dimensions to the construction of the perturbation matrix. Type normalization includes performing normalization transformation on numerical fields, performing distribution encoding or multi-hot vector expansion on enumerated fields, and using methods such as word embedding, context aggregation, or label mapping to uniformly represent text fields, thereby constructing an initial representation vector group. To improve the semantic adversarial ability, a multi-scale semantic construction mechanism can also be introduced on some fields. For example, through field segment segmentation, nested label construction, or dynamic context encoding methods, a hierarchical field expression is generated to enhance the semantic elasticity retained after dimensionality reduction perturbation.
[0037] Please refer to Figure 3 , which is a schematic flowchart of a data conversion method provided by an embodiment of the present application. Figure 3 The method shown can be applied to step S2 of the foregoing method, and the specific steps are as follows: S2.1: Extract features and encode the multiple field information data to construct an initial representation vector group; Specifically, due to the significant differences in the structural types, numerical distributions, and semantic boundaries among the original field information, even if the fields are modeled after preprocessing, it is still easy to cause problems such as feature redundancy, dimensional noise, or incomplete semantic coupling, thereby affecting the execution effect of the subsequent perturbation strategy. In this embodiment, instead of directly using the traditional language model embedding, a set of word meaning features under a controllable scale is generated through key phrase screening, stem mapping, and semantic label construction mechanisms, and is embedded in the structured representation stream in a window splicing manner.
[0038] In this embodiment, when constructing the initial representation vector group, a field-level attention filtering mechanism is also included. According to the average importance or variance distribution of candidate fields in historical training samples, fields with greater weight contribution are dynamically selected to enter the encoding process, reducing the interference of non-keyword fields on the overall semantic modeling. The vector group is finally represented in a two-dimensional tensor manner, with each row corresponding to a feature segment or semantic unit of a field, having a unified length and semantic layer classification identifier.
[0039] S2.2: Perform irreversible semantic modeling on the initial representation vector group according to a predefined perturbation strategy to generate a semantic embedding representation; The disruption strategy is used to disrupt the mapping structure between the initial representation vector group and the feature space; The purpose of this step is to construct a semantic structure representation that still has expressive power under the response task but cannot be restored to the field-level representation through the model or calculation. The disruption strategy is based on mechanisms such as random dimensionality reduction, reconstruction error feedback, and cross-coupling of multiple perturbation matrices, weakening the original semantic boundary and interrupting the field dimension path to form an irreversible vector representation. During the disruption process, the semantic embedding representation is limited to a low-dimensional, compressed, and mixed expression format, without retaining the field layer structure and label traces, but still retaining the semantic density and response gradient optimized based on the task.
[0040] The specific steps of step S2.2 of the foregoing method are as follows: S2.2.1: Perform the first round of random dimensionality reduction operation on the initial representation vector group to generate the first-round dimensionality reduction embedding vector; Specifically, to interrupt the continuous expression path of the original field information in the feature space, the first step in constructing an irreversible semantic model is to introduce a random dimensionality reduction mechanism to map the high-dimensional field representation to a low-dimensional semantic embedding space. The core purpose is to eliminate the semantic boundary and distribution inertia of the field during the original encoding process, forcing an unpredictable projection transformation of the feature expression. By reducing the dimension, the original distance structure and category distribution between fields are compressed or distorted, thereby weakening the possibility of its restoration.
[0041] In this embodiment, first, a randomly generated Gaussian matrix with a matching dimension is used as the dimensionality reduction transformation matrix; this matrix is not trained through the model and is only initialized once when constructing the semantic embedding. Its row vectors are normalized to form a set of low-rank projection bases. Subsequently, the original high-dimensional initial representation vector group is projected into the low-dimensional space using matrix multiplication to generate the first-round dimensionality reduction embedding vector. This dimensionality reduction result does not retain the field-level semantic boundary and does not depend on the preset dictionary structure, so it has initial irreversibility while retaining the overall coarse-grained semantic distribution trend.
[0042] As an example, to avoid weakening the semantic expression ability caused by information loss introduced by dimensionality reduction, the dimension of the projection matrix is set with reference to the balance relationship between the field information entropy and the structural redundancy, generally controlled between 1 / 4 and 1 / 6 of the original dimension, so as to ensure that there is still sufficient semantic carrying capacity in the compressed space.
[0043] S2.2.2: Perform reverse reconstruction calculation on the first-round dimensionality reduction embedding vector to obtain the reconstruction error distribution characteristics between the first-round dimensionality reduction embedding vector and the initial representation vector group; Specifically, the purpose of this step is to evaluate the degree of semantic information distortion introduced by random dimensionality reduction, and accordingly identify high-sensitivity regions to provide a differential strategy for the construction of the subsequent perturbation matrix. Directly performing homogeneous perturbation on all dimensionality reduction vectors will make it difficult to control the risk of semantic loss and cannot perform key perturbation on sensitive dimensions.
[0044] In this embodiment, a pseudo-inverse matrix is used to reconstruct the first-round dimensionality reduction embedding vectors. The linear back-projection is performed on the embedding vectors using the pseudo-inverse of the previous random Gaussian matrix to restore an approximate original representation. Subsequently, the Euclidean distance is calculated between the reconstructed vector and the true initial representation vector in each dimension to obtain the reconstruction error value for each dimension, and finally a complete set of error intensity vectors is constructed. By normalizing and smoothing these error vectors with a sliding window, the central tendency of the error information in each semantic dimension can be obtained.
[0045] S2.2.3: Initialize at least two perturbation matrices according to the reconstructed error distribution characteristics, where the tuning rules and tuning amounts of each perturbation matrix are different from each other and act on different feature subspaces of the first-round dimensionality reduction embedding vectors; Specifically, based on the error-dominated dimension clusters identified in the previous stage, the first-round dimensionality reduction embedding vectors are divided into multiple feature subspaces, and each subspace corresponds to a dimension cluster. The core purpose of constructing multiple perturbation matrices is to adopt differential perturbation strategies for irreversible processing for different types of feature subspaces, thereby enhancing the perturbation complexity of the overall structure in multiple directions and avoiding the formation of predictable perturbation trajectories under a unified perturbation direction.
[0046] In this embodiment, for each feature subspace, an independent perturbation matrix is constructed. The perturbation matrix contains a set of main perturbation kernels initialized along the error-dominated direction, and its structure consists of the main direction vector of the error cluster and a set of Gaussian perturbation terms. The perturbation kernel ensures the independence of each perturbation matrix in the interference strategy by controlling the perturbation norm, perturbation sparsity, and the non-linear coupling relationship between the covariance matrix and other perturbation kernels. Each perturbation matrix also sets independent tuning rules and tuning amounts, including the update strategy of the perturbation direction, the trigger condition for perturbation activation, the amplitude control of the perturbation gain, the dynamic adjustment method of the perturbation probability threshold, etc.
[0047] In an example, the specific steps for initializing the perturbation matrix are as follows: S2.2.3.1: Calculate the error intensity distribution of the reconstructed error distribution characteristics in different feature dimensions; S2.2.3.2: Divide the first-round dimensionality reduction embedding vectors into at least two feature subspaces according to the error intensity distribution, where each feature subspace corresponds to a main dimension cluster of the error intensity; S2.2.3.3: Initialize the main perturbation kernel of the perturbation matrix according to the direction vector of the main dimension cluster, where the main perturbation kernel includes perturbation basis vectors constructed along the direction vector and Gaussian perturbation terms; S2.2.3.4: Determine the tuning rule and tuning amount of the perturbation matrix according to the statistical characteristics of the error in the eigen-subspace and in combination with the structural characteristics of the main perturbation kernel.
[0048] In this embodiment, the initialization step of the perturbation matrix is used to construct multiple heterogeneous perturbation matrices for the specific feature structure of the reconstruction error distribution after the first round of random dimensionality reduction of the initial representation vector group, so as to achieve the goal of structural disruption in semantic embedding representation. Among them, the perturbation matrix is not generated through a unified template, but is partitioned and constructed after spatial decomposition of the reconstruction deviation trend of the dimensionality-reduced vector in different dimensions, and a differential response mechanism is formed within each subspace, finally realizing a nested coupling type of irreversible perturbation.
[0049] As an example, in S2.2.3.1, the distribution of the reconstruction error between the first-round dimensionality-reduced embedding vector and the original initial representation vector group is calculated. The design intention of this step is to identify which feature dimensions or dimension groups have a large degree of semantic information loss during the dimensionality reduction and compression process, so as to provide a distribution basis for subsequent perturbation operations. Although theoretically the dimensionality reduction operation can be regarded as an information-preserving transformation, in the actual modeling process, due to factors such as redundancy, collinearity or uneven semantic local density in high-dimensional vectors, the dimensionality reduction process is often accompanied by the structural decoupling of some highly sensitive dimensions, making the semantic reconstruction accuracy in some regions significantly lower than that of other dimensions. Therefore, making these dimension differences explicit and using them for perturbation strategy differentiation becomes one of the key points to improve the perturbation accuracy and irreversible ability.
[0050] In this embodiment, the calculation process of the reconstruction error is completed using the pseudo-inverse reprojection technique, that is, without introducing a learnable mapping model, the original representation vector group is approximated through the inverse mapping, and then the error value of each dimension is obtained through vector difference. In order to eliminate the influence of local noise on the error clustering result, this step also introduces operations such as sliding window smoothing, interval truncation and local statistical fitting, so as to ensure that the formed error distribution curve has continuity and clustering operability. The finally obtained error distribution vector can not only be used as the basic basis for subspace division, but also provide the upper and lower limit boundary conditions for perturbation intensity adjustment.
[0051] As an example, in S2.2.3.2, according to the obtained results of the reconstruction error distribution, the first-round dimensionality reduction embedding vectors are divided into several feature subspaces according to the error-dominant dimensions. This division is not a simple equal-slice operation, but a dynamic clustering division based on the error density distribution, so that the dimensions within each feature subspace have high similarity in terms of reconstruction performance or have similar semantic sensitivity indicators. Since the dimension mapping relationship often shows non-uniform characteristics of local compression and global sparsity during the dimensionality reduction process, directly adopting a unified perturbation strategy on the overall embedding space may instead cause some important information to be missed and perturbed, or some low-noise regions to be mis-perturbed, leading to semantic ambiguity.
[0052] In this embodiment, instead of using a unified perturbation matrix to act on the entire embedding space, at least two or more feature subspaces are divided, and exclusive perturbation response logics are designed within each subspace. Such a division method of subspaces can not only improve the pertinence of perturbation, but also avoid problems such as mutual interference and signal masking during the perturbation process in different regions, which helps to construct a local perturbation structure with stronger controllability and more independent action paths, thereby enhancing the anti-reduction ability of the overall embedding expression model.
[0053] As an example, in S2.2.3.3, for each feature subspace, a perturbation matrix is initialized separately, and its core is composed of a main perturbation kernel. The construction logic of the main perturbation kernel is generated based on the error-dominant direction corresponding to the aforementioned subspace. To ensure that the perturbation has spatial sensitivity and difference in the acting direction, in this step, the eigenvalue decomposition of the covariance structure of each subspace is performed, and the principal component direction is extracted as the basis for constructing the perturbation basis vector. This direction vector not only carries the structural trend information within the subspace, but also implies the directional semantics of the most likely reconstruction loopholes formed by this subspace during the dimensionality reduction process, so it is very suitable as the perturbation direction expansion path.
[0054] In this embodiment, during the actual construction process, the perturbation kernel is not only determined by the main direction, but also an additional Gaussian perturbation term is added to enhance the non-stability and unpredictability of the perturbation response. The Gaussian perturbation term can be set as a group of random variables with a mean of zero and a controlled variance, which is used to simulate non-linear perturbation behavior and avoid the reversible derivation of the perturbation trajectory. Through this superposition method of the directional perturbation basis and the perturbation noise term, each perturbation matrix is not only separated from other matrices in terms of the acting range, but also has a structural difference in the perturbation path, further strengthening the indecomposability of the embedded representation structure after perturbation.
[0055] As an example, in S2.2.3.4, according to the error statistical characteristics of each subspace and the construction characteristics of its corresponding perturbation kernel, an independent set of tuning rules and tuning amounts is configured for each perturbation matrix. The tuning rules include strategic mechanisms such as whether the perturbation direction can be adaptively adjusted, whether the perturbation frequency is activated based on feedback control, and whether cross-matrix linkage feedback is introduced in the perturbation response. The tuning amount involves quantifiable parameters such as the perturbation intensity boundary, the perturbation step coefficient, and the perturbation probability threshold. These configurations not only ensure the heterogeneity of each perturbation matrix when performing perturbation tasks but also enable the entire perturbation system to have dynamic response capabilities and adaptability.
[0056] In this embodiment, to maintain the heterogeneity of the control mechanism among the perturbation matrices during the action process and maximize the non-uniformity of the perturbation trajectory, the system configures an independent set of tuning rules and corresponding tuning amounts for the error information density of each eigen-subspace, the type of error distribution (whether it is concentrated, whether there are boundary fluctuations), and the direction stability and perturbation sensitivity of its main perturbation kernel, forming a dynamic adjustment mechanism for the perturbation matrix. This configuration mechanism is neither a template-based reuse nor a fixed parameter input but is driven by the statistical data collected during the previous dimensionality reduction and reconstruction stages to solve and select parameters. Its core goal is to ensure that the response behavior of each perturbation path has both response diversity and does not introduce structural conflicts or spatial couplings.
[0057] Specifically, during the configuration process of the tuning rules, the system first evaluates the response space of the perturbation direction based on the direction offset stability of the main perturbation kernel, that is: in continuous perturbation responses, if the angle between the main direction and the feedback vector oscillates frequently within a certain range, it indicates that there is a phase drift phenomenon in the perturbation direction, and the direction update mechanism needs to be enabled; at this time, the perturbation direction adaptive adjustment strategy will be activated in the tuning rules, and it is set that when the amplitude of the feedback vector exceeds the upper bound of the projection intensity of the current main perturbation direction vector, the perturbation direction reconstruction function is triggered, and the main perturbation kernel direction is continuously rotated around the axis of the feedback vector, and the rotation amplitude increases in a proportional relationship with the angle value. On the other hand, if the subspace error response shows a periodic amplification-suppression trend and the convergence speed of the perturbation response function is lower than expected, the system will determine it as a low-sensitivity and high-inertia subspace, and enable the perturbation frequency delay activation mechanism in the tuning rules, that is, trigger a new round of perturbation response every fixed number of rounds or when the error rises significantly, ensuring that the perturbation quality will not be reduced due to frequent perturbations. Further, in the subspace where there is a covariance structure among multiple perturbation matrices, if a certain subspace has been indirectly affected during non-local perturbations, the system will dynamically detect the variance mutation of its local perturbation response function and activate the cross-feedback linkage mechanism in the tuning rules, writing the second-order perturbation response from other perturbation matrices into the current perturbation basis vector through a coupling function to perform a dynamic gain rotation, thereby realizing asynchronous interference compensation between different perturbation kernels.
[0058] In terms of the setting of the tuning amount, the system first constructs a perturbation intensity reference interval based on the global intensity average value and local fluctuation amplitude of the subspace error. The upper and lower boundaries of this interval respectively correspond to the maximum error signal growth and minimum response compression amplitude after the current subspace is perturbed, thus defining the dynamic domain of the perturbation amplitude. If this subspace still maintains a low-amplitude response under multiple perturbations, the system automatically tightens its upper limit of perturbation intensity and raises the lower limit to enhance the perturbation expression density. In terms of the perturbation step size, the initial setting of the step size coefficient refers to the projection density of the main perturbation direction in the embedding space, that is, the cosine value of the angle between the perturbation basis vector and the current vector group. The closer the value is to 1, the more redundant the current perturbation is in this space, and the step size should be appropriately reduced to increase the perturbation frequency; otherwise, the step size is increased to quickly cross the direction damping area. Similarly, the perturbation probability threshold is not a static constant, but is dynamically updated based on the response sensitivity of the main perturbation kernel direction to the cross-feedback vector. Before the perturbation operation, the system will judge whether to trigger the perturbation operation through pre-perturbation exploration calculation (that is, simulating the response trend after a small projection of the current perturbation direction). If the fluctuation value of the pre-response function is lower than the set micro-perturbation threshold, the current perturbation behavior is suppressed. This mechanism ensures that perturbations will not be frequently executed in inefficient regions, thereby saving computing power resources and also improving the response expectation value of the perturbation operation.
[0059] Furthermore, to prevent the perturbation matrix from falling into a stable mode during long-term operation, the tuning mechanism is also configured with a periodic perturbation self-check module: when the perturbation increment of a certain perturbation matrix converges to zero or a near-zero interval within multiple subspace periods, the system will consider that its perturbation kernel has entered a high-stability area or encountered feedback saturation, and automatically reset its current tuning amount or re-randomize the perturbation direction, thus breaking the inertial evolution trend of the perturbation path.
[0060] It should be noted that since the tuning rules and tuning amounts of each perturbation matrix are non-shared structures, there is no risk of unified pacing or collaborative failure during the overall perturbation process. On the contrary, through the differential tuning strategy, the perturbation behaviors among the matrices show asymmetric responses and non-uniform intensity distributions, so that any perturbation trajectory in the embedding representation cannot be used to reverse-infer the overall field information, greatly enhancing the irreversible property of semantic modeling.
[0061] S2.2.4: Perturb the corresponding eigen-subspace according to the perturbation matrix, and generate a local perturbation expression according to the perturbation response, where the perturbation also has a cross-effect based on the perturbation responses of the eigen-subspaces corresponding to other perturbation matrices; Specifically, the perturbation operation is not limited to performing random perturbations in the local feature subspace, but also introduces a cross-feedback mechanism across subspaces. The purpose is to make each perturbation behavior affected by the perturbation results of other subspaces simultaneously, so as to form a dynamic perturbation coupling structure and enhance the unpredictability and irreversibility of the overall expression. The traditional piecewise perturbation method often has security risks such as structural disassemblability and perturbation response rollback due to the lack of cross-space feedback.
[0062] In this embodiment, the perturbation matrix first performs a perturbation operation on its corresponding feature subspace, including feature offset along the perturbation kernel direction, injecting Gaussian noise with a specific covariance structure, and applying a dynamic weight of perturbation intensity. Subsequently, the system collects the differences, co-variations, and relative angles between the current perturbation response vector and the perturbation response vectors of other subspaces, and constructs a cross-feedback vector for further guiding the adjustment of the local perturbation direction.
[0063] Furthermore, when each perturbation matrix executes the next perturbation cycle, it performs update actions such as direction rotation, perturbation gain scaling, or perturbation condition reset according to the calculation result of the angle between the cross-feedback vector and the current perturbation direction. The finally obtained perturbation response is no longer the result of single-direction perturbation, but a response synthesis in a multi-perturbation coupling field, constituting a set of local perturbation expressions with nesting, irreversibility, and dynamic adaptability.
[0064] In an example, the specific steps for perturbing the feature subspace are as follows: S2.2.4.1: Perform a perturbation operation on the feature embedding vectors in their respective feature subspaces according to the perturbation matrix to generate a preliminary perturbation response. The perturbation operation includes feature offset and injection of random Gaussian noise; Specifically, the structured perturbation kernel in the previous stage is applied to the actual sub-sections of the embedding vector space. Since each feature subspace aggregates a class of similar main dimensions based on dimensionality reduction and error partitioning, local perturbations are more likely to concentrate on affecting potential high-semantic-intensity regions. By means of a biased movement based on the perturbation kernel direction and superimposing a random Gaussian perturbation term with adjustable covariance, the embedding vectors can be deviated from the original stable direction to form an initial expression body after perturbation.
[0065] In this embodiment, the perturbation operation adopts the following strategy: first, the perturbation direction vector is obtained according to the main perturbation kernel in the perturbation matrix, and then a directed amplitude offset is performed in this direction. The offset can be a fixed step size or dynamically scaled according to the local intensity function of the error in the subspace. At the same time, an additional set of Gaussian noise vectors that satisfy the zero mean distribution and the dimension co-correlation structure are introduced. These noise terms act on the current feature embedding through local injection to ensure that the final perturbation expression contains both directional perturbations and unpredictable perturbations. The effect is to form a noisy perturbation response, which contains irreversible perturbation trajectories.
[0066] Furthermore, the disturbance operation can be completed in one go, or multiple rounds of superimposed disturbance can be used, wherein the disturbance intensity of each round decreases, but the disturbance direction is continuously corrected with feedback to achieve stable coverage of multi-dimensional disturbance.
[0067] S2.2.4.2: For any perturbation matrix, when performing the perturbation operation, a cross-feedback vector is generated according to the preliminary perturbation response of the corresponding characteristic subspace of the other perturbation matrices; Specifically, if all perturbation matrices operate independently, separable perturbation intervals may be formed between subspaces, which can be deconstructed by potential reverse technology. By introducing cross-feedback, the closed nature of the perturbation path can be effectively broken, so that the perturbation direction is controlled by multi-space interference input, and the unpredictability of the perturbation is improved.
[0068] In this embodiment, the cross-feedback vector is generated as follows: after each perturbation matrix generates its own preliminary perturbation response, the post-perturbation vector is differentially calculated with the pre-perturbation vector to obtain a perturbation gradient response vector. At the same time, all other perturbation matrices will also generate corresponding gradient responses. For any perturbation matrix, the perturbation direction update of its current cycle will take into account these gradient signals from other subspaces to construct a fused perturbation direction vector containing a weight adjustment factor.
[0069] Furthermore, in order to prevent all disturbance vectors from converging due to feedback merging, the covariance weighting coefficient or cosine similarity adjustment coefficient is used in the fusion process between feedback vectors to ensure that the direction update is selective. The essence of this feedback mechanism is that disturbance drives disturbance, and the disturbance behaviors of different subspaces affect each other to form a disturbance chain.
[0070] In the specific implementation, in order to construct such a perturbation chain, each perturbation matrix needs to set the perturbation feedback monitoring cycle and the cross signal fusion strategy. The monitoring cycle can be a fixed round or dynamic entropy value control, and the feedback fusion strategy can be angle driven, intensity driven or principal component driven.
[0071] S2.2.4.3: Adjust the perturbation direction of the current perturbation matrix according to the cross-feedback vector to generate a secondary perturbation response, specifically including: calculating the angle between the cross-feedback vector and the eigen-subspace direction vector; Rotate the perturbation direction of the current perturbation matrix around the cross-feedback vector by the angle to generate an updated perturbation direction; Specifically, since the perturbation direction may not be consistent with the semantic direction of the eigen-subspace, direct replacement may lead to perturbation imbalance or inability to maintain subspace coupling. Therefore, it is more reliable to adopt a vector rotation mechanism.
[0072] In this embodiment, the rotation method of the perturbation direction is as follows: First, calculate the angle between the current perturbation direction vector and the cross-feedback vector, and obtain the angle value through the vector dot product method; then, with the cross-feedback vector as the rotation axis, use a three-dimensional space vector rotation model (Rodrigues formula or quaternion rotation) to rotate the current perturbation direction around this axis by a specified angle. The rotation angle can be a full angle or scaled proportionally to achieve micro-perturbation deflection. The rotated perturbation direction will be used for the next round of perturbation operation to complete one iteration of the perturbation direction with feedback.
[0073] Furthermore, the adjustment of the rotation angle can also introduce an adaptive tuning mechanism to determine the rotation amplitude according to the variance of the cross-feedback or its orthogonality with the current perturbation direction. If the feedback vector appears repeatedly in multiple subspaces, angle suppression or perturbation redirection can be performed to avoid the perturbation resonance problem.
[0074] S2.2.5: Nest and combine the local perturbation expressions to generate a semantic embedding representation; Specifically, the local perturbation expressions of each eigen-subspace after perturbation cannot individually form a complete semantic model, and structural fusion and semantic integration are still required. The simple splicing method cannot express the correlation between cross-space perturbations. Therefore, a nested combination mechanism is needed to construct a semantic embedding representation with hierarchical, combinatorial, and context coupling capabilities.
[0075] In this embodiment, a hierarchical nested combination strategy is adopted to divide each local perturbation expression into multiple nested levels according to the perturbation intensity, structural complexity, and subspace semantic weight. The outer layer expression mainly carries the low-risk semantic response area, and the core layer aggregates the structural information of the high-intensity perturbation area. This nested structure supports subsequent response modules to perform structural unpacking or response regulation based on access requirements, and can provide nested path support for access entropy determination.
[0076] S2.3: Package the semantic embedding representation into a semantic representation model; Specifically, although the semantic embedding representation after perturbation processing already has irreversibility, it has not yet formed a complete response-capable model structure. To enable the semantic embedding representation to have the ability to respond to subsequent access, in this embodiment, it is structurally encapsulated to generate a unified semantic representation model. The encapsulation process includes semantic structure layer registration, index path binding, and access control identifier injection. In the encapsulated structure, each semantic unit or semantic fragment is mapped to a set of response dimension labels for semantic intention matching in the access request. At the same time, to avoid leaking the dimension space mapping path within the structure, this embodiment does not retain an explicit two-way mapping table from fields to semantic dimensions, but only retains a one-way index path from the response path to the original fields, and the index ontology is independent of the main model structure.
[0077] As an example, in an embodiment of the present invention, there is another method for managing and storing information of business candidate personnel. When an external access request is received, the information management and storage method further includes: S4: Parse the external access request; Specifically, since the semantic representation model of business candidate personnel does not contain traceable field information in terms of structure, it cannot be accessed through traditional field matching methods. To achieve semantic-level response generation, after the access request arrives, it is necessary to first parse its semantic intention for structural-level semantic matching in the representation model. In this embodiment, the external access request usually includes information such as the requester identifier, request task type, target position feature description, or natural language query intention. The parsing operation is achieved by constructing an access intention vector, which is generated using the same encoding specification as the semantic representation model to ensure structural comparability between the two. The construction of the access intention vector may include: context segmentation, label extraction, and structural mapping of the natural language description, and finally generating a set of nested semantic identification vectors as the input for the subsequent matching process.
[0078] S5: Match the semantic response dimension from the semantic representation model according to the parsing result; Specifically, since the semantic representation model is constructed using an irreversible perturbation strategy, the field information has been disrupted, compressed, and nested and combined, and only the responsive semantic hierarchical structure is retained. Therefore, the matching process needs to be calculated layer by layer based on the semantic similarity between the access intent vector and the semantic nested structure. In implementation, the similarity between the access intent vector and all nested nodes in the candidate personnel semantic representation model is calculated. This process constructs a multi-factor matching scoring function using indicators such as the cosine distance in the vector space, the nested path overlap rate, and the semantic center deviation degree. When multiple semantic nodes meet the similarity threshold, the candidate set can be refined according to the path level priority or the structural coupling degree, and finally a set of response dimension node sets is output. The response dimension set represents the semantic regions that can be activated in the semantic model for the current access request, and these regions will be used as the dimension index in the response generation stage to trigger the subsequent field data query path.
[0079] S6: Generate a response result according to the semantic response dimension, where the response result is queried from the original field data according to the index path; Specifically, since the semantic model and the original field information adopt a separate storage mechanism, the field data is not included in the model ontology. Only after the semantic response is triggered, the corresponding segment in the field buffer is queried through the structure index path. In this embodiment, each semantic nested node is bound with a set of field index identifiers during construction. This identifier is not directly disclosed, but can be used to drive the field query module when the semantic dimension is activated. During the response process, according to the field index path corresponding to each node in the semantic response dimension set, the field information in the restricted buffer is accessed, and the minimum coverage segment is extracted for structured output. The field granularity control policy can be set for this query process to limit security policies such as the maximum number of response fields and the maximum number of field characters, to prevent speculation of field content through multiple rounds of calls.
[0080] S7: After the response is completed, encapsulate the semantic representation model according to the access control configuration; Specifically, to prevent the access results from being iteratively extracted by multiple rounds of calls to the original model structure, it is necessary to perform dynamic encapsulation processing on the semantic representation model. In this embodiment, according to the coverage range of the semantic response, the access level of the access subject, and the access entropy value of this request, it is determined whether to trigger the encapsulation strategy and select the encapsulation method. The encapsulation strategy may include: structural compression encapsulation, that is, deleting the unactivated semantic branches and reconstructing the response hierarchy tree; summarization encapsulation, that is, converting the response path structure into a semantic label index description to hide the structure depth; perturbation enhancement encapsulation, that is, applying mild perturbation to the already responded semantic nodes to prevent the same response path from being repeatedly accessed in the next round of requests. The access entropy value is calculated based on a comprehensive evaluation of parameters such as the response path length, semantic distribution density, and access history times, and is used to dynamically determine the risk level after the response. The encapsulated semantic model will replace the original model for subsequent semantic interactions to ensure the closed-loop use and dynamic self-protection of the model structure.
[0081] The specific steps of S5 are as follows: S5.1: Construct an access intent vector; Specifically, in the technical architecture of the present invention, since the candidate personnel semantic representation model has interrupted the direct mapping relationship between fields and semantic expressions through the perturbation strategy, the visitor cannot directly retrieve data by field names or field values. Therefore, it is necessary to convert the query target content in the external access request into a semantic expression vector with the ability to compare with the semantic nested structure. This semantic expression vector is the access intent vector, and its task is to extract semantic key information from natural language or structured input and construct a comparable semantic representation using the same or compatible encoding strategy as the nested structure in the semantic model.
[0082] In this embodiment, the construction process of the access intent vector includes the following technical links: First, perform intent parsing on the original request text, use a pre-trained semantic recognition model for part-of-speech tagging, named entity recognition, and context window extraction to extract keyword groups or tags with semantic pointing; then embed the extracted information into the vector space, and the encoding method used is the same as the field embedding structure in the candidate personnel semantic model, such as using a Transformer encoder fine-tuned in the field or a dedicated semantic nesting encoder. The constructed access intent vector not only retains the semantic center of the original request but also retains the inter-word structure relationship and context information, thus providing an accurate semantic mapping basis for subsequent semantic matching.
[0083] Furthermore, in the access intent vector construction stage, a structure entropy adjustment factor can be introduced according to the requester's identity, access frequency, and organizational role to perform entropy weight adjustment on the finally generated semantic vector, so as to enhance the recognition ability of the response area in the semantic overlap area, and at the same time suppress the repeated activation of the frequently called semantic paths in the model and improve the balance of the call distribution.
[0084] S5.2: Match the access intent vector with the semantic embedding representation in the semantic representation model. The matching includes semantic similarity calculation and structural path intersection; Specifically, the candidate person semantic representation model is not generated by arranging the original fields in a fixed position, but by generating a nested structure of cross-expression under multiple subspaces through a perturbation matrix. Each semantic node represents a combination of several fields after irreversible transformation and their semantic fusion. Therefore, the matching process must be based on the similarity between vector representations in the semantic space and the path relationship formed by the nested structure in the model to accurately locate the semantic region corresponding to the access intent.
[0085] In this embodiment, the semantic matching operation first performs semantic similarity calculation, that is, compares the access intent vector with all nested nodes in the candidate person semantic model one by one. The similarity evaluation method adopts a multi-factor calculation strategy, which not only includes basic cosine similarity or Euclidean distance metrics, but also introduces structural context alignment evaluation, such as the hierarchical position of nodes in the nested structure, the path length of the parent-child relationship graph, etc., to form a composite matching scoring mechanism. To improve the fault tolerance and sensitivity of the matching, a fuzzy semantic extension strategy is also introduced. In the case where the nested distribution of some semantic words does not completely overlap, the response layer is filled by guided weight translation.
[0086] Subsequently, after completing the semantic similarity scoring, the set of candidate nodes with matching scores within the threshold is continued to be judged for structural path intersection. This path intersection refers to whether the derivation path mapped by the access intent vector has an overlapping relationship with the semantic path between the nested nodes in the model, including path prefix consistency, path nested level alignment, and path topology matching. In the case where the semantic similarity is low but the structural paths are extremely close, the matching enhancement strategy is allowed to be triggered, that is, it is regarded as an effective intersection node to improve the matching coverage rate.
[0087] S5.3: Screen out semantic nodes with semantic similarity greater than a preset threshold or path coincidence degree greater than a preset condition from the matching results; Specifically, to ensure the technical rigor of the selection of the semantic response dimension and avoid over-activation, the system needs to perform refined screening among all preliminary matching nodes to output an executable set of semantic nodes. This set constitutes the response target area of this access request, and its screening criteria should comprehensively consider two core indicators: semantic similarity and structural coincidence degree.
[0088] In this embodiment, the screening process executes the following strategy: The system preset similarity threshold and path overlap threshold, sort the similarity scores of each preliminary matching node, and directly include the nodes with scores higher than the similarity threshold in the response set. For the nodes with similarity scores lower than the similarity threshold, if their path overlap index is higher than the path overlap threshold (such as the path prefix length ratio exceeding a certain proportion, or the structural levels being completely aligned), they are also determined to have a semantically matching with structural dependency and are included in the response set.
[0089] Furthermore, the system can perform a semantic sparsification operation on the matching result set to avoid multiple nodes with similar similarities but overlapping covered fields being activated simultaneously, resulting in redundant response data. The sparsification is based on the cross-analysis of the field dimensions covered by semantic nodes, retains the nodes with the highest semantic information density, and eliminates the edge overlapping nodes.
[0090] S5.4: Use the set of index dimensions covered by the semantic nodes as the semantic response dimension; Specifically, for each nested semantic node in the candidate personnel semantic representation model, a corresponding index path mapping has been established during construction, and this mapping points to some or all of the field segments in the original field data. In the model ontology, these paths do not directly expose the field content, but during the semantic response process, they can serve as an intermediate bridge between the response logic and the field data. Therefore, using the set of index paths covered by the selected semantic nodes as the semantic response dimension is the key mechanism to achieve the compatibility of "structural response + data isolation".
[0091] In this embodiment, the system extracts the field index path identifiers recorded by each node during the model construction phase according to the selected set of semantic nodes and forms an index dimension set. This set will be used to guide subsequent field data access queries and can also be used as the structural description of the semantic response result to support the caller in understanding the semantic source of the response.
[0092] To ensure the principle of minimum available data access, the system can execute a minimum coverage optimization strategy on the index dimension set, that is, only activate the minimum subset of fields that meet the current access intention to avoid the response of invalid information. This set can be set with a lifecycle label throughout the call chain to control its validity only within the current access task and prevent the semantic dimension information from being held for a long time, resulting in the leakage of the model structure.
[0093] The access control configuration includes an access entropy adjustment strategy, and the access entropy adjustment strategy includes: S7.1: Calculate the access entropy value, where the access entropy value is calculated based on the nested level depth, the number of response dimensions, the structural path density, and the coverage ratio of the original field index path in the semantic response dimension hit by the access request; S7.2: Select an encapsulation strategy according to the access entropy value, where the encapsulation strategy includes structural compression encapsulation, summarization encapsulation, and scrambling encapsulation.
[0094] A business candidate information storage and management system, the system includes: A modeling module for obtaining field information data of business candidates and constructing an irreversible quantization representation model containing a semantic nested structure; A perturbation module for performing random dimensionality reduction, error analysis, and multi-perturbation matrix cross-control on the initial representation vector to generate an irreversible semantic embedding representation; A storage module for separately storing the semantic representation model and the original field data and establishing an index path connection structure; An access module for parsing an external access request, matching the semantic response dimension, and generating a response result through the index path; An encapsulation module for calculating the access entropy value according to the access entropy adjustment strategy and selecting a corresponding encapsulation method to perform structural compression, summarization, and scrambling operations on the semantic representation model.
[0095] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for managing and storing information of business candidate personnel, characterized in that, The information storage method described above includes: Obtain multiple field information data of the business candidate personnel, perform feature extraction and encoding on the multiple field information data, and construct an initial representation vector group; Perform irreversible semantic modeling on the initial representation vector group according to a preset perturbation strategy to generate a semantic representation model. The irreversible semantic modeling includes setting the initial representation vector group and performing a reconstruction error distribution, obtaining a feature subspace according to the main dimension cluster of the error, and initializing a perturbation matrix for each feature subspace. The tuning rules and tuning amounts of each perturbation matrix are different. The perturbation matrix acts on the corresponding feature subspace and generates a cross-feedback vector according to the perturbation response of the feature subspace. A secondary perturbation response is generated according to the cross-feedback vector to update the corresponding feature subspace, generate a semantic embedding representation, and generate a semantic representation model according to the semantic embedding representation; Store the semantic representation model in a preset information database.
2. The method for managing and storing information of business candidate personnel according to claim 1, wherein, The perturbation strategy includes: Perform a first-round random dimensionality reduction operation on the initial representation vector group to generate a first-round dimensionality reduction embedding vector; Perform reverse reconstruction calculation on the first-round dimensionality reduction embedding vector to obtain a reconstruction error distribution feature between the first-round dimensionality reduction embedding vector and the initial representation vector group; Initialize at least two perturbation matrices according to the reconstruction error distribution feature and act on different feature subspaces of the first-round dimensionality reduction embedding vector respectively; Perturb the corresponding feature subspace according to the perturbation matrix, and generate a local perturbation expression according to the perturbation response, where the perturbation also has a cross effect based on the perturbation responses of other perturbation matrix corresponding feature subspaces; Perform nested combination on the local perturbation expression to generate a semantic embedding representation.
3. The method for managing and storing information of business candidate personnel according to claim 2, wherein, Initializing at least two perturbation matrices according to the reconstruction error distribution feature includes: Calculate the error intensity distribution of the reconstruction error distribution feature on different feature dimensions; Divide the first-round dimensionality reduction embedding vector into at least two feature subspaces according to the error intensity distribution, where each feature subspace corresponds to a main dimension cluster of an error intensity; Initialize the main perturbation kernel of the perturbation matrix according to the direction vector of the main dimension cluster, where the main perturbation kernel includes a perturbation basis vector constructed along the direction vector and a Gaussian perturbation term; Determine the tuning rules and tuning amounts of the perturbation matrix according to the statistical characteristics of the error in the feature subspace and in combination with the structural characteristics of the main perturbation kernel.
4. The method for managing and storing information of business candidate personnel according to claim 3, wherein The tuning rules include a perturbation direction update method, an amplitude gain strategy, and a perturbation activation condition, and the tuning amounts include a perturbation intensity range, an update step size, and a perturbation probability threshold.
5. The method for managing and storing information of business candidate personnel according to claim 2, wherein Perturbing the corresponding feature subspace according to the perturbation matrix includes: Perform a perturbation operation on the feature embedding vectors in their respective feature subspaces according to the perturbation matrix to generate a preliminary perturbation response. The perturbation operation includes feature offset and injection of random Gaussian noise; For any perturbation matrix, when performing the perturbation operation, generate a cross-feedback vector according to the preliminary perturbation responses of other perturbation matrix corresponding feature subspaces; Adjust the perturbation direction of the current perturbation matrix according to the cross-feedback vector to generate a secondary perturbation response.
6. The method for managing and storing information of business candidate personnel according to claim 5, wherein Adjusting the perturbation direction of the current perturbation matrix according to the cross-feedback vector includes: Calculate the angle between the cross-feedback vector and the eigen-subspace direction vector; Rotate the perturbation direction of the current perturbation matrix around the cross-feedback vector by the angle to generate an updated perturbation direction.
7. The method for storing information of business candidate personnel according to claim 1, characterized in that, When an external access request is received, the information storage method further includes: Parse the external access request; Match the semantic response dimension from the semantic representation model according to the parsing result; Generate a response result according to the semantic response dimension, where the response result is queried from the original field data according to the index path; After the response is completed, encapsulate the semantic representation model according to the access control configuration.
8. The method for managing and storing information of business candidate personnel according to claim 7, characterized in that The matching the semantic response dimension from the semantic representation model according to the parsing result includes: Construct an access intention vector; Match the access intention vector with the semantic embedding representation in the semantic representation model, and the matching includes semantic similarity calculation and structural path intersection; Filter semantic nodes with semantic similarity greater than a preset threshold or path overlap greater than a preset condition from the matching results; Use the set of index dimensions covered by the semantic nodes as the semantic response dimension.
9. The method for managing and storing service candidate personnel information according to claim 7, wherein, The access control configuration includes an access entropy adjustment policy, and the access entropy adjustment policy includes: Calculate the access entropy value, where the access entropy value is calculated based on the nested level depth, the number of response dimensions, the structural path density, and the coverage ratio of the original field index path in the semantic response dimension hit by the access request; Select an encapsulation policy according to the access entropy value, where the encapsulation policy includes structural compression encapsulation, summarization encapsulation, and scrambling encapsulation.
10. A business candidate information management and storage system for implementing a business candidate information management and storage method as described in any one of claims 1-9, characterized in that, The system includes: A modeling module, configured to obtain the field information data of business candidate personnel and construct an irreversible non-quantized representation model including a semantic nested structure; A perturbation module, configured to perform random dimensionality reduction, error analysis, and multi-perturbation matrix cross control on the initial representation vector to generate an irreversible semantic embedding representation; A storage module, configured to separately store the semantic representation model and the original field data and establish an index path connection structure; An access module, configured to parse an external access request, match the semantic response dimension, and generate a response result through the index path; An encapsulation module, configured to calculate the access entropy value according to the access entropy adjustment policy and select a corresponding encapsulation method to perform structural compression, summarization, and scrambling operations on the semantic representation model.
Citation Information
Patent Citations
Text recognition method based on language model
CN120012079A
Large NLP language model privacy protection method based on differential privacy
CN116502263A
Private data protection method and device, storage medium and computer program product
CN119323051A
Human-post matching recommendation method based on BERT and latent semantic algorithm model
CN119377490A
Secure encryption method and device for parameters and messages, equipment and storage medium
CN119420564A