Office model training and data processing method, office system and storage medium
By using multimodal training data and reinforcement learning mechanisms in the vertical field, the problem of poor multimodal data fusion in the office system is solved, semantic consistency and training efficiency are improved, and business logic is adapted to vertical field.
Patent Information
- Application Number
- CN202510310739.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-11
AI Technical Summary
In vertical fields, the difficulty of efficient integration of multimodal data in office systems leads to poor semantic consistency and heavy training tasks.
By using multimodal training data in the vertical field, a dynamic selection mechanism of preset reward factors is constructed, and the reward signal is feedbacked by matching the results of the real reasoning output and the target reasoning output, reinforcement learning is carried out to iterate the office model and optimize its training process.
It improves the multimodal semantic consistency and training efficiency of the office model in the vertical field, enhances the comprehensive understanding of complex office tasks, and adapts to the subtle differences in business logic in the vertical field.
Smart Images

Figure CN120296412A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of model training, and particularly to a method for training and data processing of an office model, an office system, and a storage medium. Background Art
[0002] In the context of the rapid development of Large Language Models (LLMs), with their powerful semantic understanding and multi-task processing capabilities, they can provide a technical foundation for the implementation of intelligent office systems in vertical fields such as finance, healthcare, and law. Such office systems can parse user instructions and generate multi-modal responses. However, in high-precision scenarios in vertical fields, the practical application of models still faces problems such as heavy training tasks for office systems and poor semantic consistency caused by the difficult efficient fusion of multi-modal data. Summary of the Invention
[0003] This application provides a method for training and data processing of an office model, an office system, and a storage medium, to at least solve the problems of poor semantic consistency and heavy training tasks for office models in related technologies.
[0004] This application provides a method for training an office model, including: inputting a training sample instruction into the office model, and after the office model performs inference on the training sample instruction, obtaining the true inference output of the office model; wherein, the office model is applied to a vertical field, and the training sample instruction includes multi-modal data in the vertical field; obtaining a pre-constructed target output set, and querying the target inference output corresponding to the training sample instruction in the target output set; analyzing whether the true inference output matches the target inference output, and selecting a preset reward factor adapted to the matching result from multiple preset reward factors as a reward signal, and feeding back the reward signal to the office model to iterate the office model.
[0005] This application also provides a method for data processing, including: obtaining an instruction to be executed; using the office model to perform inference on the instruction to be executed to obtain an execution output; wherein, the office model is trained by using the method for training an office model described above.
[0006] This application also provides an office system, including: an instruction input module for inputting an instruction to be executed; an inference module connected to the instruction input module for executing the method for data processing described above.
[0007] This application also provides a computer-readable storage medium, in which a computer program is stored, and wherein the computer program, when executed by a processor, implements the steps of any one of the above methods for training an office model.
[0008] The present application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above office model training methods when executing the computer program.
[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above office model training methods when executed by a processor.
[0010] Based on multi-modal training data in a vertical domain, the present application uses multi-modal data in the vertical domain as training sample instructions to train an office model, enabling the office model to fully learn industry-specific knowledge and cross-modal association features in each vertical domain, enhancing the comprehensive understanding ability of complex office tasks. Moreover, a dynamic selection mechanism for preset reward factors can be established. By matching the results of the real inference output and the target inference output, corresponding feedback is provided to assign different preset reward factors to the office model as reward signals, providing fine-grained reinforcement learning signals for the office model, enabling the office model to be fine-tuned through different preset reward factors obtained by the reward mechanism during iteration, and improving the training efficiency. Through the process of continuously applying the reward mechanism for reinforcement learning to iterate the office model, the office model gradually adapts to the subtle differences in the business logic of the vertical domain. This method can solve the technical problems of poor semantic consistency and heavy office model training tasks through methods such as multi-modal data fusion in the vertical domain and a reward mechanism for presetting the level of reward signals during the reinforcement fine-tuning process, achieving the technical effect of ensuring multi-modal semantic consistency while reducing the training burden of the office model. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0012] Figure 1 It is an application environment diagram of a training method for an office model provided by an embodiment of the present application;
[0013] Figure 2 It is a flowchart of a training method for an office model provided by an embodiment of the present application;
[0014] Figure 3 It is another flowchart of a training method for an office model provided by an embodiment of the present application;
[0015] Figure 4 It is another flowchart of a training method for an office model provided by an embodiment of the present application;
[0016] Figure 5 Schematic diagram of a data processing method provided by an embodiment of the present application;
[0017] Figure 6 Schematic diagram of an office system structure provided by an embodiment of the present application;
[0018] Figure 7 Another schematic diagram of an office system structure provided by an embodiment of the present application;
[0019] Figure 8 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0020] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0021] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0022] It should be noted that the terms "S1", "S2", etc. are only used for the purpose of describing steps and do not particularly refer to the meaning of order or sequence. Nor are they used to limit the present application. They are only used to conveniently describe the method of the present application and should not be construed as indicating the order of steps. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.
[0023] The training method of the office model provided by the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 12 communicates with the server 14 through the network. Among them, the terminal 12 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers and portable wearable devices, and the server 14 can be implemented by an independent server or a server cluster composed of multiple servers.
[0024] To enable those skilled in the art of the present technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific implementation manners.
[0025] An embodiment of this application provides a method for training an office model. The method will be described in detail in combination with the execution process of the method for training the office model.
[0026] In one embodiment, as Figure 2 shown, Figure 2 is a schematic flowchart of a method for training an office model provided by an embodiment of this application.
[0027] S101: Input the training sample instruction into the office model. After the office model infers the training sample instruction, obtain the true inference output of the office model; wherein, the office model is applied to a vertical domain, and the training sample instruction includes multi-modal data in the vertical domain.
[0028] In this embodiment, the office model is pre-trained through the training sample instruction in the vertical domain, that is, when the training sample instruction is input into the office model, the true inference output obtained through the inference of the office model will be obtained.
[0029] Among them, pre-training generally refers to designing a language model training task based on a large-scale corpus (including language training materials such as sentences, paragraphs, etc.), training a large-scale neural network algorithm structure to learn and implement. The finally obtained large-scale neural network algorithm structure and parameters are the pre-trained language model. For subsequent other tasks, feature extraction or task fine-tuning can be performed on the basis of this model to achieve specific task purposes. The idea of pre-training is to first train a task to obtain a set of model parameters, then use this set of model parameters to initialize the network model parameters, and then use the initialized network model to train other tasks to obtain a model adapted to other tasks. By pre-training on a large-scale corpus, the neural language representation model can learn powerful language representation capabilities and can extract rich syntactic and semantic information from the text. The pre-trained language model can provide word elements (tokens) containing rich semantic information and sentence-level features for downstream tasks, and can also directly perform fine-tuning for downstream tasks on the pre-trained model to conveniently and quickly obtain a downstream-specific model.
[0030] The neural network algorithm structure of the pre-trained language model training can be CNN (Convolutional Neural Networks), RNN (Recurrent Neural Networks), LSTM (Long Short-Term Memory), etc., or it can be a model built with an attention network, such as transformer (a deep learning model), GPT (Generative Pre-trained Transformer), etc., which is not limited in this application. An attention network refers to a network model that uses an attention mechanism for training. The model assigns different weights to each part of the input sequence, thereby extracting more important feature information from the input sequence, so that the model ultimately obtains a more accurate output.
[0031] Training sample instructions are multimodal input data that guide vertical office models to complete specific tasks. They can contain professional scenario information in at least one data form, such as text, images, and tables. Its core structure consists of task instructions and business data in the corresponding field.
[0032] Vertical fields refer to sub-fields that focus on specific industries or professional scenarios, and their data and tasks are highly professional and closed, such as the financial field, medical field, legal field, etc. In other words, vertical fields are technical fields and / or enterprise types (such as hospitals, car companies, schools, etc.) to which the office mode in this embodiment is actually applied. Multimodal data refers to input or output forms that integrate multiple data forms, such as text modality, image modality, voice modality, video modality, etc.
[0033] That is to say, in the process of training the office model, the multimodal data in the vertical field to be applied in the office mode can be used as training sample instructions for training the office model. The multimodal training sample instructions are input into the office mode to optimize the semantic consistency performance of the office model in processing multimodal data. First, the vertical field data (documents, tables, charts, audio, etc.) are preprocessed by cross-modal alignment to construct annotated data sets such as image-text matching and speech-text synchronization; then, a multimodal encoder and dynamic fusion module are designed, and synchronous training is performed through contrastive learning and cross-modal attention mechanisms, while injecting domain knowledge graphs as constraints; a progressive training strategy is adopted to first independently optimize the representation capabilities of each modality, and then strengthen the semantic consistency between modalities through joint training.
[0034] S102: Obtain a pre-built target output set, and query the target inference output corresponding to the training sample instruction in the target output set.
[0035] In this embodiment, the target output set refers to a preset set of correct answers, which includes the standard output corresponding to the training sample instruction, that is, the target inference output.
[0036] During the training process of the office model, the office model can perform inference based on the training sample instructions input to it to obtain its current true inference output. To verify whether the current true inference output is reliable, the target inference output corresponding to the training sample instruction can be queried from the pre-constructed target output set.
[0037] In other words, the target inference output in the target output set is used to compare with the true inference output obtained by the office model, so as to use the target output set as a benchmark for evaluating the inference quality of the office model. The target output set can be constructed through expert annotation and business rule libraries, covering multi-modal structured outputs such as text, images, and tables, to ensure the output accuracy of the office model in the vertical domain.
[0038] S103: Analyze whether the true inference output matches the target inference output, select a preset reward factor that matches the matching result from multiple preset reward factors as the reward signal, and feedback the reward signal to the office model to iterate the office model.
[0039] In this embodiment, different levels of preset reward factors are given to the true inference output of the office model through the reward rules in the reinforcement learning stage as the reward signal. Among them, the preset reward factor can be different scores. Feedback each score to the office model, and the office model can perform fine-tuning of the office model according to the score corresponding to each output, so that the output of the office model is closer to the standard output.
[0040] Among them, reinforcement learning refers to a learning mechanism in which the model learns how to map from states to actions to maximize the obtained rewards, and continuously optimizes the state-action correspondence through the rewards given by the environment.
[0041] Fine-tuning refers to performing small-scale training on the basis of a pre-trained model for specific task objectives (downstream tasks) and task data (downstream data), achieving minor adjustments to the parameters of the pre-trained model, and finally obtaining a model adapted to specific tasks and data.
[0042] It should be noted that, different from traditional reinforcement learning algorithms, in this embodiment, during the process of reinforcement fine-tuning the office model, the training processes of heavy scoring algorithms / models, value functions, etc. can be reduced. In this embodiment, a standard output adapted to the training sample instructions, that is, the target inference output, is pre-constructed. The target inference output and the training sample instructions can also be real business data in the vertical domain at historical moments, which can reduce the acquisition burden of the training sample set. Moreover, in this embodiment, the reward signal can be associated with the matching result between the real inference output and the target inference output. Compared with complex scoring mechanisms, in this embodiment, the reinforcement fine-tuning of the office model can be completed through the significantly reduced dimensionality of the reward signal.
[0043] This embodiment is based on multi-modal training data in the vertical domain, enabling the office model to fully learn industry-specific knowledge and cross-modal correlation features in each vertical domain, and enhancing the comprehensive understanding ability of complex office tasks. A dynamic selection mechanism for preset reward factors is established, and the corresponding feedback of the result of matching the real output and the target output is fed back to different preset reward factors of the office model, providing fine-grained reinforcement learning signals for the office model. The office model obtains different preset reward factors through the reward mechanism during iteration for fine-tuning, improving the training efficiency. Through the process of continuously applying the reward mechanism for reinforcement learning to iterate the office model, the office model gradually adapts to the subtle differences in the vertical domain business logic. This method improves the professionalism, accuracy, and adaptability of the multi-modal semantic consistency office model through the reinforcement learning method of vertical domain multi-modal data fusion and the reward mechanism.
[0044] In one embodiment, as Figure 3 shown, Figure 3 is a schematic flowchart of another training method for the office model provided by the embodiment of the present application.
[0045] S201: Determine whether the real inference output of the office model conforms to the preset output format.
[0046] In this embodiment, when it is determined that the real inference output conforms to the preset output format, step S202 is executed. When it is determined that the real inference output does not conform to the preset output format, step S208 is executed.
[0047] In this embodiment, first, it is determined whether the real inference output of the office model conforms to the preset output format, and this condition is used as a prerequisite for the reward mechanism. If the real inference output of the office model does not conform to the preset output format, the fourth score is directly assigned to this output. The preset reward factors in the reward mechanism include multiple different scores, which can be the first score, the second score, the third score, and the fourth score. Specifically, how many scores are divided is not limited here. The fourth score is the lowest score among the preset reward factors that is lower than the first score, the second score, and the third score.
[0048] For example, when calculating the procurement cost, if the office model does not output according to the preset clause number structure, such as "Article X_Title_Text", even if the content is correct, it will receive a fourth score due to formatting errors, such as -2 points. Therefore, in the iteration, the formatting deviation will be corrected first to avoid system parsing failures caused by non-standard formatting, improving the output standardization and business adaptability of the office model in vertical domain applications.
[0049] In this embodiment, it is also possible to traverse the previous reward history, further determine the number of times that the true reasoning output of the office model does not conform to the preset output format, and give a lower reward signal than the previous time when the same error occurs according to the increase in the number of occurrences, so as to avoid the office model making the same mistake repeatedly through a lower score and reduce the probability of serious errors in the office model.
[0050] S202: Analyze whether the true reasoning output matches the target reasoning output.
[0051] In this embodiment, in response to determining that the true reasoning output matches the target reasoning output, step S203 is executed to evaluate the matching degree between the true reasoning output and the target reasoning output.
[0052] In response to the true reasoning output not matching the target reasoning output, step S207 is executed to use the second score as the reward signal. Among them, the second score in the preset reward factor is greater than the fourth score and lower than the first score and the third score. For example, the fourth score can be -2, the first score can be 5, the third score can be 3, and the second score can be -1; or, the fourth score can be 0 points, the first score can be 1 point, the second score can be 0 points, the third score can be 0.5 points, etc., which will not be elaborated here.
[0053] As described above, the training sample instructions in this embodiment are multi-modal data. Correspondingly, the true reasoning output and the corresponding target reasoning output in this embodiment can also be multi-modal data. The verification process of whether the reasoning outputs of each output type match will be exemplified later. The reasoning output refers to the true reasoning output and the target reasoning output, which will not be elaborated here first.
[0054] In this embodiment, by further determining whether the real inference output matches the target inference output, the comparison result is first divided into two results: match and non-match. If the two outputs do not match, a lower second score is selected from the preset reward factors as the preset reward factor. Taking whether the real inference output matches the target output as the core judgment condition of the reward mechanism, according to the matching situation between the inference output of the office model and the target output, corresponding reward signals are fed back. A higher score is given when there is a match to encourage the office model to optimize its inference ability; a lower score is given when there is no match to prompt the office model to correct the error. This reward mechanism based on the matching degree can effectively improve the accuracy and generalization ability of the office model and reduce the occurrence of incorrect inferences.
[0055] S203: Evaluate the matching degree between the real inference output and the target inference output.
[0056] In this embodiment, the matching degree between the real inference output and the target inference output is evaluated by the method of executing S204 to determine whether the matching degree is greater than a preset matching degree threshold, and no limitation is imposed on other specific evaluation dimensions here.
[0057] S204: Determine whether the matching degree is greater than a preset matching degree threshold.
[0058] In this embodiment, when it is determined that the matching degree is greater than the preset matching degree threshold, step S205 is executed to use the first score as the reward signal; when it is determined that the matching degree is not greater than the preset matching degree threshold, step S206 is executed to use the third score as the reward signal.
[0059] Among them, as described above, the preset reward factors may include a first score, a second score, and a third score, and the third score is less than the first score and greater than the second score.
[0060] In this embodiment, the matching degree between the true inference output and the target inference output is evaluated by comparing the matching degree with the matching degree threshold, that is, the reward mechanism for output matching is further refined. When it is determined that the matching degree is greater than the preset matching degree threshold, the first score is used as the reward signal; when it is determined that the matching degree is less than the preset matching degree threshold, the third score is used as the reward signal. Among them, when it is equal to the preset threshold, the first score or the third score can be used as the reward signal, and it is not limited here which score is obtained specifically when it is equal to the matching degree threshold. By introducing the third score and based on the evaluation mechanism of the matching degree threshold, the setting of the reward signal is further refined. Compared with the method that only uses two-level scoring, this solution allows for more precise quantification of the matching result and enhances the dynamic regulation ability of the inference quality of the office model. When the inference output highly matches the target output, the highest reward (the first score) is given to encourage the office model to maintain high-quality inference ability; when the matching degree is lower than the threshold but still has a certain degree of reasonableness, a medium reward (the third score) is given to avoid the learning stagnation of the office model caused by overly strict punishment; only when the matching degree is low, the lowest reward (the second score) is imposed to guide the office model to optimize. This hierarchical reward strategy can effectively improve the fine-grained learning ability of the office model, reduce misjudgment, improve the adaptability of the office model to complex and multi-modal data, make its inference results more accurate and stable, and thus optimize the intelligent performance of the office automation system.
[0061] S205: Use the first score as the reward signal.
[0062] In this embodiment, the preset reward factors include the first score and the second score, and the first score is greater than the second score.
[0063] S206: Use the third score as the reward signal.
[0064] In this embodiment, the preset reward factors further include the third score, the third score is less than the first score and greater than the second score.
[0065] S207: Use the second score as the reward signal.
[0066] S208: Use the fourth score as the reward signal.
[0067] In this embodiment, the preset reward factors further include the fourth score, and the fourth score is less than the second score.
[0068] That is, the scores of the preset reward factors are in the order of the first score, the third score, the second score, and the fourth score.
[0069] S209: Feedback the reward signal to the office model to iterate the office model.
[0070] In this embodiment, through a reward mechanism, corresponding different reward signals are given to the real inference outputs generated by the office model, and the respective reward signals are fed back to the office model. The office model fine-tunes the office model according to the received feedback signals to iterate the office model, so that the real inference output of the office model is closer to the target inference output, and finally an office model adapted to specific tasks and data is obtained.
[0071] Further, the verification process of whether the inference outputs of each output type match is exemplified below.
[0072] Analyzing whether the real inference output matches the target inference output may further include the following steps.
[0073] Specifically, the output types of the real inference output and the target inference output can be obtained; wherein, the output types include at least one of a numerical type, a text type, and an image type.
[0074] In response to the output types of the two including the numerical type, it is determined whether the output numbers of the real inference output and the target inference output are the same; in response to the output numbers of the two being the same, it is determined that the output numbers of the real inference output and the target inference output match; in response to the output numbers of the two being different, it is determined that the output numbers of the real inference output and the target inference output do not match.
[0075] In response to the output types of the two including the text type, the output texts of the real inference output and the target inference output are obtained; it is determined whether the proportion of the same texts in the output texts of the two is greater than the proportion threshold; in response to the proportion of the same texts being greater than the proportion threshold, it is determined that the output texts of the real inference output and the target inference output match; in response to the proportion of the same texts being less than the proportion threshold, it is determined that the output texts of the real inference output and the target inference output do not match.
[0076] In response to the output types of the two including the image type, the output images of the real inference and the target inference output are obtained; it is determined whether the image similarity of the output images of the two is greater than the similarity threshold; in response to the image similarity being greater than the similarity threshold, it is determined that the output images of the real inference output and the target inference output match; in response to the image similarity being less than the similarity threshold, it is determined that the output images of the real inference output and the target inference output do not match.
[0077] Specifically, if the outputs of the real inference output and the target inference output are obtained, where the real inference output is 85 and the target inference output is 85, the output type is the numerical type. Since 85 = 85, it is determined that the output numbers match. If the real inference output is 80 and the target inference output is 85, since 80 ≠ 85, it is determined that the output numbers do not match.
[0078] Specifically, if the obtained true inference output is "The patient has persistent high fever and characteristic tongue manifestations, which conform to the relevant provisions of the 'Diagnosis Guidelines for Target Diseases'.", and the target inference output is "The child has persistent fever and typical tongue symptoms, and the diagnosis criteria in the 'Diagnosis Guidelines for Target Diseases' should be referred to for judgment.", the output type is text type. Since 19 out of 28 characters in the true output match those in the target output, the calculated matching degree = 19 / 28 ≈ 67.9% (assuming the threshold ratio is 60%). Since 67.9% > 60%, it is determined that the output text matches.
[0079] Specifically, if the obtained true inference output is a pulmonary CT scan image showing typical infection characteristics (ground-glass lesions), and the target inference output is a similar pulmonary CT scan image also showing ground-glass lesion characteristics, the output type is image type. The calculated similarity is 70% (assuming the similarity threshold is 85%). Since 70% < 85%, it is determined that the output image does not match.
[0080] Among them, the methods for calculating image similarity can be PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index), MSE (Mean Squared Error), cosine similarity, etc.
[0081] Optionally, in this embodiment, business data in a vertical domain can also be constructed; among them, the business data includes at least one of a term set in the vertical domain and a knowledge graph; and the business data is written into the office model.
[0082] Specifically, constructing business data in a vertical domain can be achieved by collecting data resources in the vertical domain. The collection methods can include automatically recording various system operations through system logs, detailed tracking of user behavior through operation records, and encouraging employees to actively provide opinions through feedback forms, etc., while ensuring the security of data during collection, transmission, and storage. The data resources include internal enterprise data, employee usage feedback and business process logs, industry terms, specification texts, knowledge bases, etc., providing data support for model optimization, generating a term set including professional terms and their definitions, context usage methods, etc., or constructing a knowledge graph to structurally organize entities, relationships, attributes, etc. in the domain to enhance the knowledge reasoning ability of the office model. The constructed business data in the vertical domain is preprocessed to adapt it to the input format of the office model, and the term set and knowledge graph are embedded into the office model by means of fine-tuning or knowledge distillation, etc.
[0083] In this embodiment, through the introduction of a term set, the office model can accurately understand and generate texts that comply with domain specifications, improving professionalism. By providing semantic associations through a knowledge graph, the reasoning ability of the office model is enhanced, enabling it to better handle complex business requirements and strengthening the industry understanding ability of the office model.
[0084] In one embodiment, as Figure 4 shown, Figure 4 FIG. is a schematic flowchart of another method for training an office model provided by an embodiment of the present application.
[0085] S301: Construct a plurality of sub-office models that adapt to the office requirements of each client.
[0086] S302: Use the training sample instructions of the client to train the associated sub-office model, and use the reward signal to iterate the sub-office model; wherein, each time the sub-office model is iterated, gradient clipping is performed on the training sample instructions, and noise data is assigned to the training sample instructions to update the training sample instructions.
[0087] In this embodiment, the detailed steps of iterating the sub-office model using the reward signal can be as described in steps S101 to S103 or S201 to S209 in the foregoing text, and will not be elaborated herein.
[0088] S303: Obtain each trained sub-office model, aggregate the model parameters of the sub-office models, and form a global office model.
[0089] S304: Use the global office model as a new sub-office model, and perform training sample instruction training for each client until the global office model is trained.
[0090] In this embodiment, a collaborative mechanism of Federated Learning (FL) and Differential Privacy Optimization (DPO) is adopted. According to the office requirements of different clients, multiple sub-office models are initialized. Each office model is adapted to a specific business scenario. Each client-side uses its own data to train the model on the local server, avoiding direct data sharing or uploading, and effectively protecting data privacy. The training process includes batch (mini-batch) sampling of the training data, calculating gradients for each data sample, and then clipping the gradient values to avoid the excessive influence of certain abnormal data on the gradients. Noise (such as Laplace noise, random noise conforming to the Gaussian distribution) is added to the clipped gradients to mask the characteristics of the original data. The intensity of this noise can be adjusted according to the privacy budget to ensure privacy protection while minimizing the impact on model accuracy. Noise can also be added in the parameter aggregation stage, and secure multi-party computation is used to enhance protection. The office model parameters are updated using the noisy gradients to complete a privacy-protected gradient update process. This optimization method can train the office model without data leakage.
[0091] Among them, federated learning is a distributed machine learning framework that supports local model training on different departments or devices and uploads model parameters rather than raw data to the central server for aggregation.
[0092] Differential privacy is a method for protecting data privacy in model training. By adding noise in gradient calculation, it ensures that the model does not leak details of sensitive data during training.
[0093] In this embodiment, after each client completes local training, it uploads the updated office model parameters to the central server. The server aggregates these parameters by weighted averaging or other means to generate an updated global model. The aggregated global model is sent to each client for a new round of local training, forming an iterative loop. This mechanism can achieve distributed collaborative learning without leaking raw data, effectively supporting the continuous optimization and upgrade of the model to meet the changing needs of enterprises.
[0094] In this embodiment, the distributed architecture of federated learning ensures that data is always stored locally, avoiding the privacy leakage risk caused by the centralization of raw data. At the same time, differential privacy is used to add noise in the parameter exchange process to prevent individual information from being inferred through office model updates, ensuring that office model training can be carried out using distributed data under the premise of privacy protection, and can also realize the dynamic evolution of the office model to meet the business needs of different stages.
[0095] In one embodiment, the training method of the office model further includes introducing a multi-modal alignment loss function to ensure semantic consistency among different modalities such as text, speech, and images. The aim is to quantify the semantic deviation between different modalities and gradually reduce this deviation through an optimization process, thereby improving the performance of the office model in multi-modal interaction scenarios.
[0096] Specifically, first, feature extraction is performed on the input multi-modal data. For text data, a large language model, such as LLM, is used to convert it into a high-dimensional semantic vector; for speech data, after converting it into text through a speech recognition module, semantic vectors are extracted, and at the same time, features such as pitch and speech rate of the speech are retained as auxiliary information; for image data, a deep convolutional neural network is used to extract visual feature vectors and map them to the same semantic space as the text. The generation of these feature vectors ensures that different-modal data can be compared and aligned in a unified semantic space.
[0097] Among them, the basic form of the multi-modal alignment loss function can be expressed as:
[0098]
[0099] Among them, x i and x j respectively represent input data from different modalities (such as text and images), f i and f j are feature extraction functions for the corresponding modalities, d(f i (x i ), f j (x j )) is a distance function (such as cosine distance or Euclidean distance), and w ij is the weight between modality pairs, used to balance the importance of different-modal alignments. For example, in an office scenario, the alignment between text and images may be more critical than the alignment between speech and images, so the semantic consistency of key modality pairs can be highlighted by adjusting the weights.
[0100] In the training of the office model, the multi-modal alignment loss function is jointly optimized with other task loss functions (such as text generation loss, image generation quality loss, etc.). The training objective is to minimize the total loss.
[0101] Among them, the basic form of minimizing the total loss function can be expressed as:
[0102] L total = L task + λL align Equation 1-2
[0103] Among them, L task is the minimum total loss, and L taskis the task-specific loss (such as the cross-entropy loss for generated text), λ is a hyperparameter used to adjust the weight of the multimodal alignment loss, and L align is the multimodal alignment loss. Through the gradient descent algorithm, the office model parameters are continuously adjusted to make the feature representations of different modalities gradually approach in the semantic space. For example, when the user inputs a voice command to generate a chart, the model will optimize to ensure that the chart content is semantically consistent with the voice command.
[0104] In one embodiment, to adapt to the diversity of enterprise office scenarios, the office model combines the enterprise-specific multimodal knowledge graph to further enhance the semantic alignment effect. Among them, the knowledge graph stores enterprise-specific terms, business processes, and modality association relationships (such as the typical correspondence between "project progress" and bar charts). During the training process of the office model, the office model dynamically adjusts the weights and constraint conditions of the loss function according to the knowledge graph. For example, if an enterprise prefers to use line charts to represent time series data, the system will prioritize strengthening the alignment training between text descriptions and line charts, thereby improving the scene adaptability.
[0105] In one embodiment, as Figure 5 shown Figure 5 is a schematic flowchart of a data processing method provided by an embodiment of the present application.
[0106] S401: Obtain the instruction to be executed.
[0107] In this embodiment, the instruction to be executed represents the user's multimodal input (text, voice, image, etc.).
[0108] S402: Use the office model to reason about the instruction to be executed to obtain an execution output; among them, the office model is trained by the training method of the office model.
[0109] In this embodiment, the execution output represents the multimodal output result obtained after the user's multimodal input is reasoned by the office model.
[0110] Specifically, when the user writes a report or prepares a presentation, the input text may contain key data, conclusions, or trend descriptions. After the instruction to be executed is parsed for text information and data statistics by the office model, it will generate auxiliary charts, infographics, or illustrations to visually present the data and analysis results.
[0111] In this embodiment, the training method of the office model is as described in the previous steps S101 to S103 or S201 to S209.
[0112] In this embodiment, the user can obtain the instruction output result inferred by the office model through multimodal instruction input. For example, when the user writes a report or prepares a presentation, the input text may contain key data, conclusions, or trend descriptions. After the office system completes the parsing of text information and data statistics, it will generate auxiliary charts, infographics, or illustrations to visually present the data and analysis results.
[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0114] In one embodiment, as Figure 6 shown, Figure 6 FIG. is a schematic structural diagram of an office system provided by an embodiment of the present application. The office system may include an instruction input module 21 and an inference module 22, and the instruction input module is connected to the inference module.
[0115] The instruction input module 21 is used to obtain the instruction to be executed.
[0116] The inference module 22 is used to infer the instruction to be executed by using the office model to obtain an execution output; wherein, the office model is trained by using the training method of the office model described in any one of the foregoing.
[0117] Further, as Figure 7 shown in Figure 7 FIG. is another schematic structural diagram of an office system provided by an embodiment of the present application.
[0118] The instruction input module 21 may further include a user interaction unit 211.
[0119] The user interaction unit 211 is used to receive the multimodal input (text, voice, image, etc.) of the user, and provide a friendly user interface, such as a text input box, a voice recognition function, an image upload interface, etc., to meet the usage habits of different users. It also includes a user interface for the user to perform multimodal input, and the user interface includes a text input box, a voice recognition function, an image upload interface, etc., to meet the usage habits of different users. Text input supports instant messaging and document editing; voice input integrates voice recognition and supports multiple languages; image input supports picture upload and parses handwritten notes, charts, etc.
[0120] The inference module 22 may further include a natural language processing unit 221, a task generation and management unit 222, an enterprise data collection and feedback unit 223, a training and optimization unit 224, a system interface unit 225, and a data security and privacy protection unit 226.
[0121] The natural language processing unit 221 is used to perform semantic parsing, intent recognition, and sentiment analysis on user input using an office model.
[0122] In this embodiment, by integrating a pre-trained office model, in-depth semantic analysis, intent recognition, and sentiment analysis are performed on the input, the understanding of specific terms is improved by combining with an enterprise term library, and multi-round interaction is achieved by maintaining the context.
[0123] The task generation and management unit 222 is used to automatically generate, assign, and track tasks according to the parsing result, and support business process management.
[0124] In this embodiment, by automatically creating tasks or workflows, functions such as task assignment, reminder, tracking, and feedback are provided.
[0125] The enterprise data collection and feedback unit 223 is used to collect enterprise internal data, employee usage feedback, and business process logs, and provide data support for model optimization.
[0126] In this embodiment, various methods are used to comprehensively collect data, such as automatically recording various system operations through system logs, detailed tracking of user behavior through operation records, and encouraging employees to actively provide opinions through feedback forms. The data is cleaned, labeled, and anonymized, and a feedback entry is provided. The log system records system operations and errors. Among them, data tagging facilitates model training, and anonymization protects privacy to ensure the security of data during collection, transmission, and storage.
[0127] The training and optimization unit 224 is used to continuously optimize the office model by adopting differential privacy optimization and federated learning technologies.
[0128] In this embodiment, the training and optimization process is as described in S301 to S304 in the previous text, and will not be elaborated here.
[0129] The system interface unit 225 is used to provide a standardized API interface (Application Programming Interface), and strongly supports data transmission, model update, and seamless integration with third-party systems.
[0130] In this embodiment, an interface is built based on the API, and communication protocols such as the Hypertext Transfer Security Protocol are used to ensure the encryption of data transmission. Relevant authentication mechanisms are used to ensure the legality and security of identities. At the same time, scalability is fully considered in the interface design to meet the needs of future business development.
[0131] The data security and privacy protection unit 226 is used to prevent information leakage, protect data by encrypting storage, implement fine-grained permission management based on role-based access control, and regularly perform security audits and vulnerability scans to ensure system security.
[0132] In this embodiment, the user's instruction input passes through the office that has undergone adaptive evolution and successively passes through the user interaction unit, the natural language processing unit, the task generation and management unit, and obtains comprehensive data feedback through the enterprise data collection and feedback unit. Then, the office model is continuously optimized in the training and optimization unit by using differential privacy optimization and federated learning technologies. Finally, the optimized office model is updated through the system update and deployment module, thus completing the entire adaptive evolution process. The architecture diagram of this office system has high flexibility and scalability, can continuously optimize and upgrade according to the changing needs of the enterprise and the dynamic changes of the external environment, realize adaptive evolution, and is conducive to keeping pace with the enterprise's business development.
[0133] Specifically, for example, Xiaoming, an enterprise employee, needs to arrange a project discussion meeting involving personnel from multiple departments. Xiaoming inputs via voice: "Help me arrange a project discussion meeting at 2 pm next Tuesday and invite colleagues from the R & D department and the marketing department to attend." After parsing the instruction, the office model identifies the time as 2 pm next Tuesday, the event as a project discussion meeting, and the participants as colleagues from the R & D department and the marketing department. The office model automatically generates a task, creates a meeting event in the enterprise calendar, reserves a suitable meeting room, and sends out meeting invitations. In terms of task management, the system tracks the response status, reminds those who have not replied, and sends a reminder notice before the meeting. Subsequently, the office model collects the task completion status and user feedback for office model training to optimize the accuracy of the next task generation.
[0134] Specifically, for example, Li Li, an employee, has completed a project report and hopes to file it and notify relevant personnel. Li Li inputs an instruction via text: "File the project report I just completed in the project repository and notify team members to view it." After parsing the instruction, the office model identifies the action as filing the report, the location as the project repository, and the notification object as team members. The system automatically executes file archiving, uploads the report to the specified location, sets the access permission so that only team members can view it, and sends them a viewing notice. In task management, the office model records the viewing situation of team members and collects feedback. The data collected is used for office model optimization to improve the processing efficiency for similar instructions.
[0135] Specifically, for example, after the office system has been running for a period of time, it has collected a large amount of user interaction data and feedback, including instruction input, task execution status, and employee feedback, covering multiple departments and business scenarios. After preprocessing, the data is cleaned of invalid or duplicate information, classified, labeled, and anonymized to protect privacy. Next, during model training, differential privacy optimization is used to add noise to the gradient calculation to protect data privacy, and at the same time, training is carried out through federated learning on the local servers of each department, and only the model parameters are uploaded to the server for aggregation. The accuracy of the office model in specific business scenarios has been improved, and user satisfaction has also increased accordingly. Finally, after security review, the office model is deployed and updated during non-working hours to ensure that business continuity is not affected.
[0136] The office system in this embodiment has significant advantages compared with traditional office systems: in terms of functional adaptability, through the coordination of multiple modules, especially the cooperation of the enterprise data collection and feedback module and the model training and optimization module, it can continuously optimize the functions of the tool according to the unique business processes and changing needs of the enterprise. Whether it is a small startup or a large multinational enterprise, it can customize and optimize the tool in a data-driven manner according to its own business characteristics, which is difficult to achieve by traditional office systems. Traditional office systems are usually standardized products and are difficult to quickly adapt to the diverse needs of different enterprises or different development stages of the same enterprise.
[0137] In terms of data security and privacy protection, the combination of differential privacy optimization and federated learning technology has built a solid defense line for data security. Against the backdrop of frequent data leakage incidents today, enterprises attach great importance to data security. In the process of data processing by traditional office systems, especially when involving multi-party data sharing and collaboration, it is difficult to effectively guarantee the privacy and security of data. However, this tool has taken strict security measures in all aspects from data collection, transmission, storage to model training to ensure that enterprise data and employee privacy are fully protected.
[0138] In terms of user experience, the multi-modal interaction method has greatly improved the usability of the system. Different users have different preferences for interaction methods. Most traditional office systems only support a single text interaction method, which limits the choices of users. However, this tool provides multiple interaction methods such as text, voice, and image, allowing users to freely switch according to the actual scenario and their own habits, improving the convenience and efficiency of operation.
[0139] In terms of improving office efficiency, the intelligent task automation function can quickly and accurately parse user instructions and generate tasks, reasonably allocate resources, and track progress. This reduces the cumbersome processes and errors of manual operations, especially when dealing with a large number of tasks, it can significantly improve work efficiency. Traditional office systems often require users to perform multiple manual operations in task management, which is time-consuming and laborious and prone to omissions.
[0140] In this embodiment, the application scenarios of the office system are not limited to the enterprise office environment, but can also extend to industries such as healthcare, education, and finance. For example, in the healthcare industry, the system can be applied to the intelligent office system of hospitals to assist doctors in medical record management, diagnosis suggestions, and patient communication, improving the efficiency and quality of medical services; in the education field, the system can be used for teaching management in schools to assist teachers in lesson preparation, homework grading, and communication with parents, enhancing the intelligent level of education and teaching; in the financial industry, the system can be applied to the office systems of banks or securities companies to assist in risk assessment, customer management, and compliance review, optimizing the financial service process.
[0141] Secondly, in terms of technology integration, the office system can combine blockchain technology to achieve data traceability and immutability, improving the credibility and security of data; introduce augmented reality technology to provide more intuitive and interactive operation and display methods in the office scenario, enhancing the user experience and work efficiency; integrate with Internet of Things technology to achieve linkage with intelligent devices, providing functions such as intelligent conference room management and device status monitoring, and creating a smart office environment.
[0142] In terms of function upgrade, the office system can strengthen the research on emotional intelligence, enhance the ability to recognize and respond to user emotions, and provide more user-friendly and considerate services; combine big data analysis technology to provide business decision-making support functions, such as market trend analysis, risk warning, and strategic planning, to assist enterprises in scientific decision-making; establish an enterprise knowledge base to support the accumulation, sharing, and reuse of knowledge, promoting enterprise internal knowledge management and innovation ability improvement.
[0143] Finally, in terms of business model innovation, the office system can provide software as a service, reducing the deployment and maintenance costs of enterprises and facilitating the use of small and medium-sized enterprises; provide customized office system functions and services according to different industries and enterprise scales to meet diverse market demands; cooperate with other software and service providers to build a complete enterprise digital ecosystem, realizing resource sharing and complementary advantages, and jointly exploring a broader market space.
[0144] Each module in the above office system can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.
[0145] The embodiment of the present application also provides an electronic device, which can be a terminal, and its internal structure diagram can be as Figure 8 shown Figure 8A schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes an image generation method. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0146] Those skilled in the art can understand that Figure 7 and Figure 8 the structure shown in
[0147] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0148] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.
[0149] Those skilled in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0150] It can further be appreciated that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0151] The above has introduced in detail a method for training an office model provided by the present application. Specific examples are used herein to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A training method for an office model, characterized in that, The training method includes: Inputting the training sample instructions into the office model, and the office model infers the training sample instructions to obtain the true inference output of the office model; wherein, the office model is applied to a vertical domain, and the training sample instructions include multi-modal data of the vertical domain; Obtaining a pre-constructed target output set, and querying the target inference output corresponding to the training sample instructions in the target output set; Analyzing whether the true inference output matches the target inference output, selecting a preset reward factor adapted to the matching result from multiple preset reward factors as a reward signal, and feeding back the reward signal to the office model to iterate the office model.
2. The training method according to claim 1, wherein The preset reward factors include a first score and a second score, and the first score is greater than the second score; The selecting a preset reward factor adapted to the matching result from multiple preset reward factors as a reward signal includes: In response to the true inference output matching the target inference output, taking the first score as the reward signal; In response to the true inference output not matching the target inference output, taking the second score as the reward signal.
3. The training method according to claim 2, wherein The preset reward factors further include a third score; the third score is less than the first score and greater than the second score; The taking the first score as the reward signal in response to the true inference output matching the target inference output includes: Evaluating the matching degree between the true inference output and the target inference output; Judging whether the matching degree is greater than a preset matching degree threshold; In response to the matching degree being greater than the matching degree threshold, taking the first score as the reward signal; In response to the matching degree being less than the matching degree threshold, taking the third score as the reward signal.
4. The training method according to claim 2, wherein, The preset reward factors further include a fourth score, and the fourth score is less than the second score; The selecting a preset reward factor adapted to the matching result from multiple preset reward factors as a reward signal further includes: Judging whether the true inference output of the office model conforms to a preset output format; In response to the true inference output conforming to the preset output format, selecting a preset reward factor between the values of the first score and the second score as the reward signal; In response to the true inference output not conforming to the preset output format, taking the fourth score as the reward signal.
5. The training method according to claim 1, wherein The analyzing whether the true inference output matches the target inference output further includes: Obtaining the output types of the true inference output and the target inference output; wherein, the output types include at least one of a digital type, a text type, and an image type; In response to the output types of the two including the digital type, judging whether the output numbers of the true inference output and the target inference output are the same; in response to the output numbers of the two being the same, determining that the output numbers of the true inference output and the target inference output match; in response to the output numbers of the two being different, determining that the output numbers of the true inference output and the target inference output do not match; In response to the output types of both being text types, obtain the output texts of the true inference output and the target inference output; determine whether the proportion of the same texts in the output texts of both is greater than a proportion threshold; in response to the proportion of the same texts being greater than the proportion threshold, determine that the output texts of the true inference output and the target inference output match; in response to the proportion of the same texts being less than the proportion threshold, determine that the output texts of the true inference output and the target inference output do not match. In response to the output types of both including image types, obtain the output images of the true inference and the target inference output; determine whether the image similarity of the output images of both is greater than a similarity threshold; in response to the image similarity being greater than the similarity threshold, determine that the output images of the true inference output and the target inference output match; in response to the image similarity being less than the similarity threshold, determine that the output images of the true inference output and the target inference output do not match.
6. The training method according to claim 1, characterized in that The feedback of the reward signal to the office model to iterate the office model includes: Construct a plurality of sub-office models adapted to the office requirements of each client; Use the training sample instructions of the client to train the associated sub-office model, and use the reward signal to iterate the sub-office model; wherein, each time the sub-office model is iterated, gradient clipping is performed on the training sample instructions, and noise data is given to the training sample instructions to update the training sample instructions; Obtain each trained sub-office model, aggregate the model parameters of the sub-office models to form a global office model; use the global office model as a new sub-office model, and respectively perform training with the training sample instructions of the client until the global office model is trained.
7. The training method according to claim 1, wherein The training method further includes: constructing the business data of the vertical domain; wherein, the business data includes at least one of the term set of the vertical domain and the knowledge graph; Write the business data into the office model.
8. A data processing method, characterized in that, The data processing method includes: Obtain the instruction to be executed; Use the office model to perform inference on the instruction to be executed to obtain an execution output; wherein, the office model is trained by using the training method of the office model according to any one of claims 1 to 7.
9. An office system, characterized in that, The office system includes: An instruction input module for inputting an instruction to be executed; An inference module connected to the instruction input module for executing the data processing method according to claim 8.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program, when executed by a processor, implements the steps of the training method of the office model according to any one of claims 1 to 7.
Citation Information
Cited By
Consistency enhancement method and device based on multi-source reward fusion, equipment and medium
CN121052330A