Data simulation method, device, electronic equipment, computer program and storage medium

CN121562230BActive Publication Date: 2026-09-22SHENZHEN ZHICHENG SOFTWARE TECH SERVICE CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610087111.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-22
Publication Date
2026-09-22
Estimated Expiration
2046-01-22

AI Technical Summary

Technical Problem

然而,脱敏后的数据通常丧失了其原始结构和统计特性,无法用于有效的软件开发、测试或机器学习模型训练,仅能用于数据展示

Benefits of technology

[0010]本申请提供的数据仿真方法、装置、电子设备、计算机程序和存储介质,通过采用基于结构相似度损失函数训练的仿真模型,将数据仿真的核心目标由传统的语义相似性转变为语句结构相似性,使得模型在训练过程中被明确引导至学习并复现输入语句的语法结构与实体类型序列,而非其具体语义内容;这一机制不仅有效剥离了生成数据与原始敏感数据在语义上的直接关联,保障了数据安全,而且由于结构相似度损失函数持续优化模型使其输出与输入在语句结构上高度一致,确保了仿真数据在语法规则和数据结构层面与真实数据保持等效性,从而能够直接替代真实数据用于模型训练和系统测试,最终实现对结构相似数据的高效仿真产出。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121562230B_ABST
    Figure CN121562230B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a data simulation method and device, electronic equipment, computer program and storage medium. Embodiments of the present application obtain a to-be-simulated sentence, input the to-be-simulated sentence into a pre-trained simulation model to output a simulated sentence similar in sentence structure to the to-be-simulated sentence. The pre-trained simulation model is trained based on at least a structural similarity loss function. The structural similarity loss function is used to guide the simulation model to make the sentence structure similarity between the output content of the simulation model and the input content of the simulation model higher than a preset threshold during training. Embodiments of the present application can ensure that the simulated data is equivalent to the real data in terms of grammar rules and data structure, so that the simulated data can directly replace the real data for model training and system testing, and efficient simulation output of the structural similar data is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data security, specifically to a data simulation method, apparatus, electronic device, computer program, and storage medium. Background Technology

[0002] In the field of data security, the protection of sensitive data is of paramount importance. Existing technologies mainly employ data anonymization methods to mask or replace sensitive data. However, anonymized data typically loses its original structure and statistical properties, making it unusable for effective software development, testing, or machine learning model training; it can only be used for data visualization.

[0003] Another similar technique is similar text generation, such as the SimBERT and RoFormer-Sim models, which aim to generate semantically similar sentences. However, this type of technique focuses on semantic similarity rather than data structure similarity. In practical applications, such as database simulation testing, what is needed is fake data with similar grammatical structures (such as similar part-of-speech sequences and entity type sequences) but different content from real data, to ensure that the test environment can realistically simulate the production environment without revealing any real information.

[0004] Therefore, how to achieve efficient simulation output of structurally similar data is an urgent problem to be solved. Summary of the Invention

[0005] This application provides a data simulation method, apparatus, electronic device, computer program, and storage medium that can achieve efficient simulation output of structurally similar data.

[0006] This application provides a data simulation method, including: Obtain the statement to be simulated; The statement to be simulated is input into a pre-trained simulation model to output a simulated statement with a similar statement structure to the statement to be simulated. The pre-trained simulation model is trained based on at least a structural similarity loss function, which guides the simulation model during training to ensure that the sentence structure similarity between the output content and the input content of the simulation model is higher than a preset threshold.

[0007] This application also provides a data simulation device, including: The target acquisition module is used to acquire the statement to be simulated. The statement simulation module is used to input the statement to be simulated into a pre-trained simulation model to output a simulated statement with a similar statement structure to the statement to be simulated; wherein, the pre-trained simulation model is trained based on at least a structural similarity loss function, which is used to guide the simulation model during training so that the statement structure similarity between the output content of the simulation model and the input content of the simulation model is higher than a preset threshold.

[0008] This application also provides an electronic device, including a memory storing multiple instructions; the processor loads instructions from the memory to execute steps in any of the data simulation methods provided in this application.

[0009] This application also provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute steps in any of the data simulation methods provided in this application.

[0010] The data simulation method, apparatus, electronic device, computer program, and storage medium provided in this application, by employing a simulation model trained based on a structural similarity loss function, shifts the core objective of data simulation from traditional semantic similarity to sentence structure similarity. This ensures that the model is explicitly guided during training to learn and reproduce the grammatical structure and entity type sequence of the input sentence, rather than its specific semantic content. This mechanism not only effectively separates the direct semantic association between the generated data and the original sensitive data, ensuring data security, but also ensures that the output and input are highly consistent in sentence structure due to the continuous optimization of the model by the structural similarity loss function. This ensures that the simulated data remains equivalent to real data at the level of grammatical rules and data structure, thus directly replacing real data for model training and system testing, ultimately achieving efficient simulation output of structurally similar data. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart illustrating the data simulation method provided in an embodiment of this application; Figure 2 This is a schematic diagram of a structural similarity matrix provided in an exemplary embodiment of this application; Figure 3 This is a schematic diagram of the structure of the data simulation device provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0014] This application provides a data simulation method, apparatus, electronic device, computer program, and storage medium.

[0015] Specifically, the data simulation device can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet, smart Bluetooth device, laptop, or personal computer (PC); the server can be a single server or a server cluster consisting of multiple servers.

[0016] In some embodiments, the data simulation device can also be integrated into multiple electronic devices, such as multiple servers, with multiple servers implementing the data simulation method of this application.

[0017] In some embodiments, the terminal may also be used as a server to perform some or all of the functions of a server.

[0018] For example, refer to Figure 1 The electronic device can be a terminal that can acquire the statement to be simulated; input the statement to be simulated into a pre-trained simulation model to output a simulated statement with a similar statement structure to the statement to be simulated; wherein the pre-trained simulation model is trained based at least on a structural similarity loss function, which is used to guide the simulation model during training so that the statement structure similarity between the output content of the simulation model and the input content of the simulation model is higher than a preset threshold.

[0019] The following sections provide detailed descriptions of each example. It should be noted that the sequence numbers of the following embodiments are not intended to limit the preferred order of the embodiments.

[0020] In this embodiment, a data simulation method is provided, such as... Figure 1 As shown, the specific process of this data simulation method can be as follows: Step 110: Obtain the statement to be simulated.

[0021] The statement to be simulated refers to the original input statement that needs to generate simulation data, and it is the reference object that needs to be processed in the simulation.

[0022] In the entire data simulation process, the "statement to be simulated" plays the role of the input source. It is the starting point for method execution and the original data that is to be protected and used to generate its "replacement". The purpose of the entire simulation is to generate a new statement based on this statement, which is highly similar in syntax but different in content.

[0023] Step 120: Input the statement to be simulated into the pre-trained simulation model to output a simulated statement with a similar statement structure to the statement to be simulated.

[0024] In this embodiment, the pre-training process of the simulation model may include the following steps: (1) Training data construction.

[0025] In this embodiment, a certain number of sample data pairs can be prepared in advance. Each sample data pair consists of a first sample statement and a second sample statement whose statement structure similarity is higher than the preset threshold.

[0026] For example, a sample data pair could be "Baiyun Mountain main peak elevation 382 meters, Baiyun Mountain height 382 meters", where the first sample statement is "Baiyun Mountain main peak elevation 382 meters" and the second sample statement is "Baiyun Mountain height 382 meters". The order of the first and second sample statements is not specified here; the former can be considered the first sample statement and the latter the second sample statement.

[0027] Understandably, when constructing sample data for training simulation models, the core task is to create "structurally similar pairs of statements of the same type." This involves two key points: "same type" and "structural similarity."

[0028] In this context, "same type" refers to sentences that belong to the same category in terms of semantics or function. This means that the paired sentences should describe the same or highly related things, answer the same questions, or perform the same linguistic actions. Taking the example above, both sample sentences describe address categories, thus belonging to the same semantic category.

[0029] "Structural similarity" includes the following two aspects, a and b.

[0030] a. Similar part-of-speech tagging sequences This means that the grammatical function categories of words in a sentence are similar. After a sentence is segmented and tagged with parts of speech, the order of the part-of-speech tags should be roughly consistent.

[0031] Taking address data as an example, the part-of-speech sequence of sentence A (Baiyun Mountain main peak is 382 meters above sea level) is: noun / noun / noun / verb / numeral / quantifier; the part-of-speech sequence of sentence B (Baiyun Mountain is 382 meters high) is: noun / adjective / numeral / quantifier.

[0032] It can be seen that although the sequence lengths differ, their core structure is "subject (nominal component) + core word describing the attribute (verb / adjective) + numerical measure (numeral + quantifier)". This correspondence of core grammatical functions forms the basis for structural similarity.

[0033] b. Similarity of named entity sequences That is, the sequence of named entity types identified in the sentence is similar, which ensures that the "entity type" framework of the things described in the sentence is consistent.

[0034] Taking address data as an example: The entity sequence of sentence A is: LOC (Baiyunshan) / NUM (382); the entity sequence of sentence B is: LOC (Baiyunshan) / NUM (382).

[0035] As can be seen, both contain entities of the types "geographic location" and "specific numerical value", and the sequence of entity types matches.

[0036] It is important to emphasize that the "structural similarity" of this invention is fundamentally different from the common "semantic similarity." Semantic similarity requires two sentences to have the same or very similar meanings (such as "The weather is very nice today" and "The sun is shining brightly today"). However, the purpose of this embodiment is not to generate sentences with the same semantics, but to generate sentences with the same grammatical framework but different specific vocabulary, thereby achieving data simulation.

[0037] In this embodiment, the sentence structure similarity can be evaluated in the manner described above. A preset threshold can be set in advance, that is, if the sentence structure similarity is higher than this preset threshold, the two sentences can be considered to be "structurally similar".

[0038] It should be noted that traditional similar sentence generation models concatenate a sample pair (AB) in opposite directions to create two training data sets (AB and BA). To avoid this repeated concatenation leading to insufficient generalization ability and monotonous output when generating structurally similar sentences, in this embodiment, each sample pair is concatenated only once (i.e., only AB is used). Furthermore, to enhance the model's subsequent learning of structural diversity, multiple sentences C, D, ... with similar structures can be introduced to construct additional training sample pairs (e.g., AC), meaning that the same input sentence A corresponds to multiple structurally similar output sentences B, C, D, ..., thereby enriching the training data.

[0039] (2) Pre-training to obtain the simulation model 2.1 Generation Loss In this embodiment, the first sample statement can be input into the initial model to guide the initial model to generate output content corresponding to the second sample statement.

[0040] In this step, the goal is to train an initial model to generate an expected output (the content corresponding to the second sample statement) based on an input (the first sample statement). This can be similar to the process by which humans learn expressions from examples when learning a language.

[0041] The initial model learns to map one sequence (the first sample statement) to another sequence (the second sample statement). In this process, the "first sample statement" is the condition and the "second sample statement" is the learning target.

[0042] In some implementations, the initial model can be a Simbert or RoFormer-Sim model, but this embodiment does not limit it.

[0043] To achieve the above mapping, the initial model employs a specific input format and attention mechanism, specifically: First, concatenate the first sample statement (A) and the second sample statement (B) into a long sequence, in the format [CLS]A[SEP]B[SEP]. Here, [CLS] is the start symbol and [SEP] is the separator.

[0044] Then, for the portion A[SEP] in the sequence (i.e., the “first half of the sentence”), the initial model uses a bidirectional attention mechanism. This means that the initial model can see all the word segments in sentence A simultaneously, thus gaining a comprehensive understanding of its structure and semantics and encoding it into a rich context vector.

[0045] For the portion B[SEP] in the sequence (i.e., the "second half of the sentence"), the model uses a one-way attention mechanism (or causal masking). This means that when predicting the i-th word of sentence B, the model can only see all the information from sentence A and the first to the (i-1)th words in sentence B. This forces the model to learn to predict the next word regressively based on the context (sentence A) and the already generated content.

[0046] As mentioned earlier, the initial model's task is to predict each position in sentence B word by word. For example, when the input is [CLS] Baiyun Mountain's main peak is 382 meters above sea level [SEP], the model needs to predict step by step that the next word is "Baiyun Mountain", then "high", then "382", and finally "meters".

[0047] After generating the output content, the loss value of the generation loss function can be determined based on the difference between the second sample statement and the output content. This step aims to create a clear signal to the initial model about the magnitude of the gap between its generated content and the desired target, thereby guiding the adjustment of model parameters.

[0048] In some implementations, for each position in sentence B, the initial model computes a probability distribution representing how likely it considers the "next word" to be each word in the vocabulary. This probability distribution is compared to a true distribution (i.e., a "one-hot encoding," where 1 is the position of the correct word and 0 is the position elsewhere). The difference between these two distributions is then measured using a cross-entropy loss function. The greater the difference, the higher the loss value.

[0049] Then, by summing or averaging the cross-entropy losses at all positions in sentence B, we can obtain the generation loss value for the entire sequence.

[0050] As an example, suppose the second sample statement B is "Baiyun Mountain is 382 meters high". At the first prediction position, the model should output the word "Baiyun Mountain". If the probability distribution given by the model shows a high probability for "Baiyun Mountain", the loss here is small; if it incorrectly gives a high probability for "Beijing", the loss here will be large. Applying this process to each word, "high", "382", and "meter", and finally summing all the losses, we get the loss value of the generative loss function.

[0051] 2.2 Structural Similarity Loss In this embodiment, not only is it required that the generated content be consistent with the target sentence at the lexical level, but more importantly, the generated output content must maintain a high degree of syntactic similarity to the first sample sentence. The structural similarity loss function is designed precisely to guide the model to learn and satisfy this structural constraint.

[0052] The calculation of the structural similarity loss function is a multi-step process. Its core idea is to map the text from the "lexical space" to the "structural space" before measuring the similarity. The specific process is as follows: 2.2.1 Structural Representation of Text Part-of-speech tagging and named entity recognition are performed on the first sample statement (A) and the output content (C) generated by the model (C differs from the real second sample statement B in the early stages of training), respectively, and they are represented as corresponding structure vectors. Specifically: First, part-of-speech tagging is performed, which segments the sentence into words and labels the part of speech of each word, such as noun, verb, adjective, etc. Then, named entity recognition is performed, which identifies named entities (such as names of people, places, organizations, etc.) in the sentence and replaces them with their common type tags.

[0053] Thus, a sentence composed of specific words is transformed into a sequence of part-of-speech tags and entity type tags, i.e., a structure vector. This process strips away the specific semantic content, retaining only the grammatical skeleton of the sentence.

[0054] As an example, sentence A: “The main peak of Baiyun Mountain is 382 meters above sea level” could be represented as LOC, noun, noun, verb, NUM, quantifier; sentence C: “The height of Xishan is 500 meters” could be represented as LOC, noun, verb, NUM, quantifier.

[0055] 2.2.2 Similarity Calculation Understandably, model training is typically performed in batches. Within a batch, statements A and C are processed as described above to obtain their corresponding structure vectors. Subsequently, the pairwise distances between all structure vectors within that batch are calculated to obtain the structure similarity matrix, for example... Figure 2 As shown.

[0056] In some implementations, the similarity metric can be cosine similarity or Hamming distance, etc. After calculation, a structural similarity matrix of size [batch size, batch size] is obtained. Each element S_ij in this matrix represents the degree of similarity between the i-th sentence and the j-th sentence in the batch in the structural space.

[0057] 2.2.3 Loss Calculation After obtaining the structural similarity matrix, it is transformed into a supervision signal for a contrastive learning task.

[0058] First, a Softmax operation is performed on each row of the similarity matrix, converting the similarity values ​​of each row in the structural similarity matrix into a probability distribution, resulting in a probability matrix. This ensures that in each row, sentences that are structurally more similar to the current sentence have a higher probability value.

[0059] Then, in the obtained probability matrix, the position corresponding to the current sentence itself (i.e., the main diagonal elements of the matrix) is masked. Since the similarity between the sentence and itself is a constant maximum value, it does not participate in the comparative learning.

[0060] Finally, the masked probability matrix is ​​treated as a prediction distribution, and a target distribution is constructed: in this row, the position label of the sentence that forms a valid sample pair with the current sentence (i.e., the positive sample) is 1, and the labels of the other positions (negative samples) are 0. Then, the cross-entropy loss between the two is calculated as the structural similarity loss value.

[0061] 2.3 Iterative Model Training After obtaining the generation loss and structural similarity loss respectively, the initial model training process enters the iterative optimization phase. During the iteration, by fusing the two losses, the parameter update direction of the initial model is guided and corrected, so that it eventually converges to an ideal state: it can accurately generate fluent text, and ensure that the generated text is highly similar to the input text in grammatical structure.

[0062] First, the two losses are weighted and summed according to preset weights to form the overall objective function of the model optimization, i.e., the total loss.

[0063] Then, the gradient of this loss value with respect to all trainable parameters of the model (such as weights and biases) is calculated using the backpropagation algorithm. The gradient indicates the direction and magnitude in which each parameter should be adjusted to reduce the loss. Specifically, the total loss can be propagated layer by layer from the output layer through backpropagation, and the gradient of the parameters at each layer can be calculated using the chain rule. The model parameters are then iteratively updated using gradient descent optimization algorithms (such as Adam, SGD, etc.) according to the calculated gradients and the preset learning rate.

[0064] In summary, the steps in 2.1-2.3 are repeated for multiple training cycles until a preset stopping condition is met. The stopping condition can be loss convergence, meaning the total loss no longer decreases significantly on a preset validation set or tends to stabilize; it can also be reaching a preset maximum number of iterations; or it can be that the sentence structure similarity between the model's output and the simulation model's input is higher than a preset threshold.

[0065] After the loop iteration is completed, a pre-trained simulation model can be obtained. The simulation model has learned the stable mapping rules from input sentences to sentences with similar structures.

[0066] The sentence to be simulated is input into a pre-trained simulation model. During the inference phase, the simulation model can use the mapping rules learned during training to parse the grammatical structure of the sentence (similar to processing the "first half of the sentence" during training). Then, based on its internal parameters, it generates the output sequence word by word in an autoregressive manner. Unlike during training, there is no real "second sample sentence" as a target behind the model at this time; instead, it generates the sequence entirely based on the probability distribution it has learned.

[0067] Therefore, the output of the simulation model is the required simulation statement. These statements are highly similar in structure to the input statement to be simulated, but the specific content is different. For example, the input statement to be simulated could be "No. 1 Jianguomenwai Avenue, Chaoyang District, Beijing", and the output simulation statement could be "No. 108 Tiyu East Road, Tianhe District, Guangzhou". The two statements have highly similar structures, but the actual content is different.

[0068] In practical applications, database tables can be read for simulation. Data with structures similar to different fields in the database can be output according to requirements. All simulation statements generated by the model are organized according to their correspondence with the original data records and written to one or more specified simulation data tables. The structure of the simulation data tables (such as table schema, field definitions, and relational constraints) is consistent with or highly corresponding to the original data tables to be simulated, thus ensuring the substitutability of the simulation data at the application level.

[0069] Ultimately, the simulation data tables stored in the database can be securely read and accessed by other downstream software systems (such as data analysis platforms, software development and testing environments, and machine learning training systems) through standard database access interfaces (such as JDBC and ODBC) or application programming interfaces (APIs).

[0070] As can be seen from the above, the embodiments of this application, by employing a simulation model trained based on the structural similarity loss function, transform the core objective of data simulation from traditional semantic similarity to sentence structure similarity. This allows the model to be explicitly guided during training to learn and reproduce the grammatical structure and entity type sequence of the input sentence, rather than its specific semantic content. This mechanism not only effectively separates the direct semantic association between the generated data and the original sensitive data, ensuring data security, but also ensures that the output of the model is highly consistent with the input in terms of sentence structure due to the continuous optimization of the structural similarity loss function. This ensures that the simulated data remains equivalent to the real data at the level of grammatical rules and data structure, thus directly replacing the real data for model training and system testing, ultimately achieving efficient simulation output of structurally similar data.

[0071] To better implement the above methods, this application also provides a data simulation device, which can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer; the server can be a single server or a server cluster composed of multiple servers.

[0072] For example, in this embodiment, the method of this application embodiment will be described in detail by taking the data simulation device specifically integrated into the terminal as an example.

[0073] For example, such as Figure 3 As shown, the data simulation device may include a target acquisition module 210 and a statement simulation module 220, as follows: (a) Target acquisition module 210, used to acquire the statement to be simulated.

[0074] (ii) A statement simulation module 220 is used to input the statement to be simulated into a pre-trained simulation model to output a simulated statement with a similar statement structure to the statement to be simulated; wherein the pre-trained simulation model is trained based on at least a structural similarity loss function, the structural similarity loss function is used to guide the simulation model during training so that the statement structure similarity between the output content of the simulation model and the input content of the simulation model is higher than a preset threshold.

[0075] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.

[0076] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0077] As can be seen from the above, the data simulation device in this embodiment obtains the statement to be simulated; inputs the statement to be simulated into a pre-trained simulation model to output a simulated statement with a similar statement structure to the statement to be simulated; wherein, the pre-trained simulation model is trained based at least on a structural similarity loss function, which is used to guide the simulation model during training so that the statement structure similarity between the output content of the simulation model and the input content of the simulation model is higher than a preset threshold.

[0078] Therefore, the embodiments of this application can ensure that the simulation data is equivalent to the real data at the level of syntax rules and data structure, so that it can directly replace the real data for model training and system testing, and achieve efficient simulation output of structurally similar data.

[0079] This application also provides an electronic device, which can be a terminal, a server, or other similar device. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers, etc.

[0080] In some embodiments, the data simulation device can also be integrated into multiple electronic devices, such as multiple servers, with multiple servers implementing the data simulation method of this application.

[0081] In this embodiment, the electronic device will be described in detail as a terminal, for example, such as... Figure 4 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically: The electronic device may include components such as a processor 310 with one or more processing cores, a memory 320 with one or more computer-readable storage media, a power supply 330, an input module 340, and a communication module 350. Those skilled in the art will understand that... Figure 3 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 310 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 320, and by calling data stored in the memory 320, it performs various functions and processes data, thereby performing overall detection of the electronic device. In some embodiments, the processor 310 may include one or more processing cores; in some embodiments, the processor 310 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 310.

[0082] The memory 320 can be used to store software programs and modules. The processor 310 executes various functional applications and data processing by running the software programs and modules stored in the memory 320. The memory 320 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 320 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 320 may also include a memory controller to provide the processor 310 with access to the memory 320.

[0083] The electronic device also includes a power supply 330 that supplies power to the various components. In some embodiments, the power supply 330 can be logically connected to the processor 310 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 330 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0084] The electronic device may also include an input module 340, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0085] The electronic device may also include a communication module 350. In some embodiments, the communication module 350 may include a wireless module, through which the electronic device can perform short-range wireless transmission, thereby providing users with wireless broadband internet access. For example, the communication module 350 can be used to help users send and receive emails, browse web pages, and access streaming media.

[0086] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 310 in the electronic device loads the executable files corresponding to the processes of one or more application programs into the memory 320 according to the following instructions, and the processor 310 runs the application programs stored in the memory 320 to realize various functions, as follows: This application's embodiments employ a simulation model trained based on a structural similarity loss function, shifting the core objective of data simulation from traditional semantic similarity to sentence structure similarity. This explicitly guides the model during training to learn and reproduce the grammatical structure and entity type sequence of the input sentence, rather than its specific semantic content. This mechanism not only effectively separates the direct semantic association between the generated data and the original sensitive data, ensuring data security, but also ensures that the output of the model is highly consistent with the input in sentence structure due to the continuous optimization of the structural similarity loss function. This guarantees that the simulated data remains equivalent to real data at the level of grammatical rules and data structure, thus directly replacing real data for model training and system testing, ultimately achieving efficient simulation output of structurally similar data.

[0087] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0088] As can be seen from the above, the embodiments of this application can ensure that the simulation data is equivalent to the real data at the level of syntax rules and data structure, so that it can directly replace the real data for model training and system testing, and achieve efficient simulation output of structurally similar data.

[0089] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0090] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the data simulation methods provided in embodiments of this application. For example, the instructions can execute the following steps: Obtain the statement to be simulated; input the statement to be simulated into a pre-trained simulation model to output a simulated statement with a similar statement structure to the statement to be simulated; wherein, the pre-trained simulation model is trained based on at least a structural similarity loss function, which is used to guide the simulation model during training so that the statement structure similarity between the output content of the simulation model and the input content of the simulation model is higher than a preset threshold.

[0091] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0092] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the methods provided in the various optional implementations of the above embodiments.

[0093] Since the instructions stored in the storage medium can execute the steps in any of the data simulation methods provided in the embodiments of this application, the beneficial effects that any of the data simulation methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.

[0094] The foregoing has provided a detailed description of a data simulation method, apparatus, electronic device, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A data simulation method, characterized in that, include: Obtain the statement to be simulated; The statement to be simulated is input into a pre-trained simulation model to output a simulated statement with a similar statement structure to the statement to be simulated. The pre-trained simulation model is trained based on at least a structural similarity loss function, which guides the simulation model during training to ensure that the sentence structure similarity between the output content and the input content of the simulation model is higher than a preset threshold. The training steps for the simulation model include: Obtain sample statement pairs, which consist of a first sample statement and a second sample statement with a statement structure similarity higher than the preset threshold. The first sample statement and the second sample statement are structurally similar statements of the same type. The same type indicates that the two belong to the same category in terms of semantic scope or function. The structural similarity indicates that the two have similar part-of-speech tagging sequences and similar named entity sequences. The first sample statement is input into the initial model, which then guides the initial model to generate output content corresponding to the second sample statement. Based on the difference between the second sample statement and the output content, the loss value of the generation loss function is determined; The first sample statement and the output content are subjected to part-of-speech tagging and named entity recognition to represent them as corresponding structure vectors; Calculate the distance between the two structural vectors to obtain the structural similarity matrix; The loss value of the structural similarity loss function is determined based on the structural similarity matrix; Based on the loss values ​​of the structural similarity loss function and the generation loss function, the parameters of the initial model are iteratively updated until the total loss of the structural similarity loss function and the generation loss function converges, thus obtaining the pre-trained simulation model.

2. The data simulation method as described in claim 1, characterized in that, The calculation of the distance between the two structural vectors to obtain the structural similarity matrix includes: Calculate the similarity metric between the two structure vectors, wherein the similarity metric is cosine similarity or Hamming distance; The structural similarity matrix is ​​constructed based on the similarity metric, and the size of the structural similarity matrix is ​​the same as the number of sample sentences.

3. The data simulation method as described in claim 1, characterized in that, Determining the loss value of the structural similarity loss function based on the structural similarity matrix includes: The similarity values ​​of each row in the structural similarity matrix are converted into a probability distribution to obtain a probability matrix; The diagonal elements of the probability matrix are masked. Based on the probability matrix after masking, the contrastive learning cross-entropy loss is calculated and used as the loss value of the structural similarity loss function.

4. A data simulation device, characterized in that, include: The target acquisition module is used to acquire the statement to be simulated; A statement simulation module is used to input the statement to be simulated into a pre-trained simulation model to output a simulated statement with a similar statement structure to the statement to be simulated. The pre-trained simulation model is trained at least based on a structural similarity loss function, which guides the simulation model during training to ensure that the statement structure similarity between the output content and the input content of the simulation model is higher than a preset threshold. The training steps of the simulation model include: obtaining sample statement pairs, each sample statement pair consisting of a first sample statement and a second sample statement with a statement structure similarity higher than the preset threshold. The first sample statement and the second sample statement are structurally similar statements of the same type, where "same type" indicates that they belong to the same category in semantics or function, and "structural similarity" indicates that their part-of-speech tags are similar. The process involves: 1) Analyzing sequence similarity and named entity sequence similarity; 2) Inputting the first sample statement into the initial model to guide the initial model in generating output content corresponding to the second sample statement; 3) Determining the loss value of the generation loss function based on the difference between the second sample statement and the output content; 4) Performing part-of-speech tagging and named entity recognition on the first sample statement and the output content to represent them as corresponding structure vectors; 5) Calculating the distance between the two structure vectors to obtain a structure similarity matrix; 6) Determining the loss value of the structure similarity loss function based on the structure similarity matrix; 7) Iteratively updating the parameters of the initial model based on the loss value of the structure similarity loss function and the loss value of the generation loss function until the total loss of the structure similarity loss function and the generation loss function converges, thus obtaining the pre-trained simulation model.

5. An electronic device, characterized in that, It includes a processor and a memory, the memory storing multiple instructions; the processor loads instructions from the memory to perform the steps in the data simulation method as described in any one of claims 1 to 3.

6. A computer program product, characterized in that, It includes a computer program / instruction that, when executed by a processor, implements the steps of the data simulation method according to any one of claims 1 to 3.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the data simulation method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Statement retelling method and statement retelling device

    CN114154481A

  • Large language model training method and device, training data construction method and device, equipment and medium

    CN118014011A

  • Video content understanding method and device based on structured grammar information, electronic equipment and storage medium

    CN120976832A

  • Quality controlled paraphrase generation

    US20240386188A1