Method and device for constructing corpus, refrigerator and computer readable storage medium
By fine-tuning the large model of the smart refrigerator and enhancing the selection of corpus variants, the corpus is automatically constructed, solving the problem of low corpus construction efficiency and improving the accuracy of voice interaction functions and user experience.
Patent Information
- Application Number
- CN202410533666.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-29
- Publication Date
- 2025-11-04
AI Technical Summary
In existing technologies, the construction efficiency of the intelligent refrigerator corpus is low, and the user participation is high, resulting in inaccurate and untimely updates of the voice interaction function and a poor user experience.
By fine-tuning a large model, generating corpus variants, and filtering and/or enhancing them, a corpus is built using similarity ranking, reducing user involvement and improving the efficiency of automated construction.
It improved the quality and update speed of the corpus, reduced labor costs, and enhanced the accuracy and user experience of the refrigerator's voice interaction function.
Smart Images

Figure CN120892569A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information processing, for example to a method and device for constructing a corpus, a refrigerator, and a computer readable storage medium. BACKGROUND
[0002] At present, with the rapid development of Internet of Things and artificial intelligence technology, smart home devices, especially smart refrigerators, have become an important part of modern families. Smart refrigerators not only provide basic refrigeration functions, but also integrate food management, health and nutrition recommendations, and interactive entertainment and other intelligent functions. These functions are often realized through user interaction with the refrigerator, and voice interaction has become one of the main interaction methods due to its convenience. The implementation of voice interaction functions requires a comprehensive corpus, however, manually writing the corpus is very time-consuming, and as the number of interaction scenarios increases, the workload of manually maintaining and updating the corpus increases dramatically.
[0003] The related technology discloses a smart home linkage question and answer data collection method, including: obtaining the state parameters of the smart home device; sending the state parameters of the smart home device to the cloud device, and receiving the first type of question sent by the cloud device, the first type of question has an associated relationship with the target scene, and the target scene is determined by the cloud device according to the state parameters of the smart home device; push the first type of question to the user and collect the user's reply statement, and establish a first corpus; send the first corpus to the cloud device, so that the cloud device trains and updates the question and answer model according to the first corpus.
[0004] In the process of implementing the embodiments of the present disclosure, it is found that at least the following problems exist in the related technology:
[0005] The related technology automatically constructs a corpus by interacting with the user to collect the user's reply statement, which improves the efficiency of constructing the corpus to some extent. However, in actual application, constructing the corpus through user interaction requires the user to directly participate in the collection process of answering questions, which may cause user boredom, and when a large amount of data is needed for model training, relying on user answers to collect corpus may be time-consuming, thereby reducing the efficiency of constructing the corpus, the refrigerator corpus cannot be updated in time and the quality of the corpus is low, resulting in inaccurate voice interaction function and poor user experience.
[0006] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present application, and therefore can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY
[0007] The following presents a simplified summary of some aspects of the disclosed embodiments in order to provide a basic understanding of such embodiments. This summary is not an extensive overview of the embodiments described in detail in the following detailed description, and is intended neither to identify key or critical elements nor to delineate the scope of such embodiments, but to present some aspects of these embodiments in a simplified form to facilitate understanding of the detailed description.
[0008] The embodiments of the present disclosure provide a method and device for constructing a corpus, a refrigerator and a computer readable storage medium, to reduce the process of user participation in constructing the corpus, improve the efficiency of automatic construction of the corpus, ensure timely updating of the corpus, improve the accuracy of the refrigerator voice interaction function, and improve user experience on the basis of ensuring the quality of corpus data.
[0009] In some embodiments, the method comprises: fine-tuning a large model according to interaction corpus associated with the functions of the smart refrigerator to obtain a target large model; inputting initial corpus into the target large model to generate corpus variants, and screening and / or enhancing the corpus variants to obtain optimized corpus; and performing similarity sorting on the optimized corpus to screen out target corpus to construct a corpus.
[0010] Optionally, fine-tuning the large model according to the interaction corpus associated with the functions of the smart refrigerator to obtain the target large model comprises: collecting basic interaction corpus associated with the functions of the smart refrigerator; performing data preprocessing on the basic interaction corpus to obtain first interaction corpus; and fine-tuning the large model according to the first interaction corpus to obtain the target large model.
[0011] Optionally, fine-tuning the large model according to the first interaction corpus to obtain the target large model comprises: labeling the first interaction corpus to obtain second interaction corpus; setting fine-tuning parameters for the selected large model; and fine-tuning the selected large model using the second interaction corpus according to the set fine-tuning parameters.
[0012] Optionally, inputting the initial corpus into the target large model to generate corpus variants comprises: inputting the initial corpus into the target large model to generate initial corpus variants; inputting the initial corpus variants into the target large model to generate new corpus variants; in the case that the new corpus variants do not meet the requirements, inputting the new corpus variants into the target large model again to generate new initial corpus variants; or in the case that the new corpus variants meet the requirements, determining the new corpus variants as corpus variants.
[0013] Optionally, screening the corpus variants comprises: labeling the training corpus as positive samples and negative samples; training a filtering model using the labeled corpus to obtain a target filtering model; and filtering the corpus variants through the target filtering model to obtain optimized corpus.
[0014] Optionally, the corpus variant is enhanced, including: enhancing the corpus variant by a simple data enhancement method; wherein the simple data enhancement method includes one or more of synonym replacement, random insertion, random exchange, random deletion, and adding stop words.
[0015] Optionally, the optimized corpus is similarity sorted to filter out the target corpus to construct the corpus library, including: similarity sorting the optimized corpus by a cosine similarity sorting technology to filter out the target corpus meeting a preset similarity condition; and constructing the corpus library based on the target corpus.
[0016] In some embodiments, the apparatus includes a processor and a memory storing program instructions, the processor is configured to execute the above-mentioned method of constructing a corpus when executing the above-mentioned program instructions.
[0017] In some embodiments, the refrigerator includes: a refrigerator body; the above-mentioned apparatus of constructing a corpus is installed on the refrigerator body.
[0018] In some embodiments, the computer-readable storage medium stores program instructions, the program instructions execute the above-mentioned method of constructing a corpus when running.
[0019] The method and apparatus for constructing a corpus, the refrigerator, and the computer-readable storage medium provided by the embodiments of the present disclosure can achieve the following technical effects:
[0020] The large model is fine-tuned according to the interactive corpus associated with the function of the intelligent refrigerator to obtain a target large model, the initial corpus is input into the target large model to generate a corpus variant, the corpus variant is filtered and / or enhanced to obtain an optimized corpus, and finally the optimized corpus is similarity sorted to filter out a target corpus to construct a corpus library. By fine-tuning the large model for generating the corpus and filtering and / or enhancing the generated corpus, the quality of the corpus library corpus is ensured, on this basis, the process of user participation in constructing the corpus library is reduced by automatically generating the corpus by the large model, the artificial cost of constructing and updating the corpus library is reduced, the efficiency of automatic construction of the corpus library is improved, the corpus library is updated in time, and thus the accuracy of the refrigerator voice interaction function is improved and the user experience is improved.
[0021] The general description above and the following description below are exemplary and explanatory only and are not intended to be limiting. BRIEF DESCRIPTION OF DRAWINGS
[0022] One or more embodiments are illustrated by way of example in the figures that are not intended to be limiting of the embodiments. Like numbers refer to like elements throughout the drawings, which are not necessarily to scale, with:
[0023] Figure 1 is a schematic diagram of a method for constructing a corpus provided by an embodiment of the present disclosure;
[0024] Figure 2 is a schematic diagram of another method for constructing a corpus provided by an embodiment of the present disclosure;
[0025] Figure 3 is a schematic diagram of another method for constructing a corpus provided by an embodiment of the present disclosure;
[0026] Figure 4 is a schematic diagram of another method for constructing a corpus provided by an embodiment of the present disclosure;
[0027] Figure 5 is a schematic diagram of an apparatus for constructing a corpus provided by an embodiment of the present disclosure;
[0028] Figure 6 is a schematic diagram of a refrigerator provided by an embodiment of the present disclosure.
[0029] Reference signs:
[0030] 800: an apparatus for cold water machine control; 801: a processor; 802: a memory; 803: a communication interface; 804: a bus; 900: a refrigerator. DETAILED DESCRIPTION
[0031] In order to enable a more detailed understanding of the features and technical content of the embodiments of the present disclosure, the implementation of the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, which are only used for reference and do not limit the embodiments of the present disclosure. In the following technical description, in order to facilitate explanation, a plurality of details are provided to provide a full understanding of the disclosed embodiments. However, one or more embodiments can still be implemented without these details. In other cases, well-known structures and devices can be simplified to facilitate the drawings.
[0032] The terms "first", "second", and the like in the specification and claims of the embodiments of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion.
[0033] Unless otherwise specified, the term "a plurality of" means two or more.
[0034] In the embodiments of the present disclosure, the character " / " represents an "or" relationship between the objects before and after it. For example, A / B means A or B.
[0035] The term "and / or" is a description of an association relationship between objects, which means that there can be three relationships. For example, A and / or B means that there are three relationships of A or B, or A and B.
[0036] The term "corresponding" can refer to an association relationship or a binding relationship. A corresponds to B means that there is an association relationship or a binding relationship between A and B.
[0037] The embodiments of the present disclosure disclose a refrigerator, which comprises a model fine-tuning module, a corpus generation module and a determination module. The model fine-tuning module is configured to fine-tune a large model according to interaction corpus associated with the function of the intelligent refrigerator to obtain a target large model. The corpus generation module is configured to input the initial corpus into the target large model to generate corpus variants, and screen and / or enhance the corpus variants to obtain optimized corpus. The determination module is configured to sort the similarity of the optimized corpus, and screen out target corpus to construct a corpus library. The refrigerator further comprises a control module, which comprises a processor electrically connected with the above-mentioned modules and configured to control the above-mentioned modules to act.
[0038] Figures 1 to 4 It is a schematic diagram of the method for constructing a corpus provided by the embodiments of the present disclosure, any one of the following methods can be executed in the refrigerator, or in a server or terminal device in communication connection with the refrigerator. In the embodiments of the present disclosure, the processor of the refrigerator is taken as the execution subject to make the description of the scheme.
[0039] Based on the structure of the above-mentioned refrigerator, as shown in Figure 1 The embodiments of the present disclosure provide a method for constructing a corpus, which comprises the following steps:
[0040] S01, the processor fine-tunes a large model according to interaction corpus associated with the function of the intelligent refrigerator to obtain a target large model.
[0041] In this step, the large model can be any open source large model, for example, the large model can be a ChatGLM (Conversational Generative Language Model) model or a Qwen (Qianwen) model, etc. In actual application process, it is necessary to collect basic interaction corpus for the function of the intelligent refrigerator for preprocessing, including denoising, standardization, etc. The basic interaction corpus needs to ensure that the corpus data covers diversified interaction scenarios for the function of the intelligent refrigerator, such as adjusting temperature, querying food material state and / or recipe recommendation, etc. For the function of opening ice making of the refrigerator, the basic interaction corpus includes "open ice making", "start ice making mode", "please help me make ice", etc. For the function of adjusting temperature of the refrigerator, the basic interaction corpus includes "adjust temperature up", "adjust temperature down", "turn off temperature adjustment", etc. Then the basic interaction corpus can be enhanced, and the enhancement content includes positive samples and negative samples, wherein the positive sample represents a sample containing valid interaction content, and the negative sample represents a sample including invalid or misleading interaction content. Specifically, taking the function of opening ice making of the refrigerator as an example, the positive sample of "open ice making" is "please help me open ice making", "I want to open ice making mode", "start ice making function", etc.; and the negative sample is "ice making ice cream", "close ice making", "is the ice making mode turned on", etc. Taking the function of recommending recipes of the refrigerator as an example, the positive sample of "recommend recipes" is "help me recommend several dishes", "tell me several delicious dishes", etc.; and the negative sample is "help me recommend several fun", "I don't want to eat dishes", etc.
[0042] S02, the processor inputs the initial corpus into the target large model to generate corpus variants, and screens and / or enhances the corpus variants to obtain optimized corpus.
[0043] S03, the processor sorts the optimized corpus according to similarity, and screens out target corpus to construct a corpus library.
[0044] By adopting the method for constructing a corpus library provided in the embodiments of the present disclosure, the large model is fine-tuned according to the interaction corpus associated with the function of the intelligent refrigerator to obtain a target large model, the initial corpus is input into the target large model to generate corpus variants, and the corpus variants are screened and / or enhanced to obtain optimized corpus, and finally the optimized corpus is sorted according to similarity, and the target corpus is screened out to construct a corpus library. By fine-tuning the large model for generating corpus and screening and / or enhancing the generated corpus, the quality of the corpus library corpus is ensured, and on this basis, by automatically generating corpus by the large model, the process of user participation in constructing the corpus library is reduced, the artificial cost of corpus library construction and updating is reduced, the efficiency of automatic construction of the corpus library is improved, the corpus library is ensured to be updated in time, thereby improving the accuracy of the refrigerator voice interaction function and improving the user experience.
[0045] Based on the structure of the refrigerator as described above, as shown in Figure 2 The embodiment of the present disclosure provides a method for constructing a corpus, comprising:
[0046] S21, the processor collects basic interaction corpus associated with the function of the smart refrigerator.
[0047] S22, the processor performs data preprocessing on the basic interaction corpus to obtain the first interaction corpus.
[0048] In this step, the data preprocessing of the basic interaction corpus includes data cleaning of the basic interaction corpus. Specifically, irrelevant characters, incorrect information and / or repeated content of the basic interaction corpus can be removed by a denoising algorithm. For example, the positive sample and negative sample set contains multiple repeated corpora, then the processor only retains one basic interaction corpus with the highest priority. The priority can be determined according to the quality of the corpus, which can be determined by the length of the corpus, the frequency of use of the corpus, the relevance of the corpus to the function of the refrigerator, etc.
[0049] S23, the processor fine-tunes the large model according to the first interaction corpus to obtain the target large model.
[0050] S02, the processor inputs the initial corpus into the target large model to generate corpus variants, and filters and / or enhances the corpus variants to obtain optimized corpora.
[0051] S03, the processor sorts the optimized corpora according to similarity, and filters out the target corpus to construct the corpus.
[0052] The method for constructing a corpus provided by the embodiment of the present disclosure collects basic interaction corpus associated with the function of the smart refrigerator, and performs data preprocessing on the basic interaction corpus to obtain the first interaction corpus. Through the data preprocessing steps such as denoising and standardization, the data quality of the first interaction corpus used for training is ensured, and more accurate and clean learning samples are provided for the large model. Finally, the processor fine-tunes the large model according to the first interaction corpus to obtain the target large model. The first interaction corpus associated with the function of the smart refrigerator is used to fine-tune the large model, so that the model can learn the language expression and user intent in the function field of the smart refrigerator, thereby improving the accuracy of the refrigerator function voice interaction and improving the user experience.
[0053] Optionally, the processor fine-tunes the large model according to the first interaction corpus to obtain the target large model, comprising: the processor labels the first interaction corpus to obtain the second interaction corpus; the processor sets fine-tuning parameters for the selected large model; and the processor fine-tunes the selected large model using the second interaction corpus according to the set fine-tuning parameters.
[0054] The fine-tuning parameters include a learning rate, a batch size, a training round, and the like. In actual application, in order to make the large model better adapt to the interactive field of the function of the smart refrigerator, a method of fine-tuning the large model can be used. In order to make the selected large model understand and generate text in a specific format, the first interactive corpus needs to be labeled. Specifically, the processor labels the first interactive corpus to obtain the second interactive corpus, including: the processor adds specific slot symbols < and > before and after the sentences of the first interactive corpus to obtain the labeled second interactive corpus. In this way, the text to be generated can be distinguished from other information. For example, the first interactive corpus is "you are a text data augmentation robot, and the seed sentence is: open the ice maker, help me generalize 5 semantically consistent sentences", and the second interactive corpus is "< please help me open the ice maker >", "", "< start the ice maker function >", and the like. The first interactive corpus is "you are a text data augmentation robot, and the seed sentence is: recommend a recipe, help me generalize 5 semantically consistent sentences", and the second interactive corpus is "< help me recommend a few dishes >", "< tell me a few delicious dishes >", and the like. When fine-tuning the large model, one or more fine-tuning parameters such as a learning rate, a batch size, and a training round are set. The preprocessed and labeled second interactive corpus is used to fine-tune the selected large model according to the set fine-tuning parameters, so that the large model learns specific expression modes and reply modes in the interactive context of the smart refrigerator through fine-tuning.
[0055] In this way, the processor labels the first interactive corpus to obtain the second interactive corpus, and through labeling the first interactive corpus, the large model can distinguish the corpus text to be generated from other information. Then, the processor sets fine-tuning parameters for the selected large model, and finally uses the second interactive corpus to fine-tune the selected large model according to the set fine-tuning parameters. By setting appropriate fine-tuning parameters, the model fine-tuning process can be optimized to prevent overfitting or underfitting.
[0056] Based on the structure of the refrigerator as shown in Figure 3 The method for constructing a corpus provided by the embodiment of the present disclosure comprises the following steps.
[0057] S01, the processor fine-tunes a large model according to interactive corpus associated with the function of the smart refrigerator to obtain a target large model.
[0058] S31, the processor inputs the initial corpus into the target large model to generate an initial corpus variant.
[0059] S32, the processor inputs the initial corpus variant into the target large model to generate a new corpus variant.
[0060] S33, in the case that the new corpus variant does not meet the requirements, the processor inputs the new corpus variant to the target large model again to generate a new initial corpus variant; or in the case that the new corpus variant meets the requirements, the processor determines the new corpus variant as the corpus variant.
[0061] In this step, the generated new corpus variant meeting the requirements includes meeting the format requirements and quality requirements. For example, the format of the new corpus variant meets the default format, and the quality of the new corpus variant meets the filtering requirements of the filtering model.
[0062] In this step, the fine-tuned large model, i.e. the target large model, can accept an initial corpus to generate one or more corpus variants, which provide different expression modes while keeping the original meaning. Specifically, optimization can be performed through multiple rounds of iteration, i.e. using the newly generated corpus variant as the seed sentence for the next round of generation, i.e. the initial corpus, so as to further expand the corpus and improve the construction efficiency of the corpus. For example, the initial corpus variant is "turn off ice making", and the new corpus variant is "<please help me turn off ice making>", "", "<stop the ice making function>", "<turn off the ice making for me>", "<stop making ice cream>", etc. The corpus variants meeting the requirements are matched and extracted by using regular expressions, and specifically, the slot symbols < and > are recognized to extract the corpus variants meeting the format requirements of the intelligent refrigerator interaction from the new corpus variants generated by the target large model. In addition, the performance of the fine-tuned target large model can be evaluated through interaction tests with professionals and real users, and the target large model can be further optimized according to the feedback, including adjusting the fine-tuning strategy, increasing or decreasing the training data, etc., so as to improve the naturalness and accuracy of the refrigerator interaction function.
[0063] S34, the processor screens and / or enhances the corpus variants to obtain optimized corpus.
[0064] S03, the processor sorts the optimized corpus according to the similarity and screens out the target corpus to construct the corpus.
[0065] Among them, the corpus variant refers to a sentence with the same meaning as the seed sentence, i.e. the initial corpus, but different expression.
[0066] The method for constructing a corpus provided by the embodiment of the present disclosure is used, the processor inputs the initial corpus into the target large model to generate an initial corpus variant, and inputs at least one initial corpus variant into the target large model to generate at least one new corpus variant, the above steps are repeated for multiple rounds of iterations until at least one corpus variant meeting the requirements is generated. The fine-tuned target large model is suitable for voice interaction of specific functions of the intelligent refrigerator, has strong language understanding and generation capabilities, and can generate multiple initial corpus variants according to the given initial corpus. Since the model capabilities can be further tapped in each round of iteration, the target large model learns from the corpus variants generated in each round, and can continuously improve the quality and applicability of the generated corpus variants.
[0067] Optionally, the processor performs screening on the corpus variants, including: the processor labels the training corpus as positive samples and negative samples; the processor trains a filtering model using the labeled corpus to obtain a target filtering model; and the processor filters the corpus variants through the target filtering model to obtain optimized corpus.
[0068] The filtering model can be any language model that can be used for filtering, for example, a BERT (Bidirectional Encoder Representations from Transformers) model, a GPT (Generative Pre-trained Transformer) model, or an ERNIE (Enhanced Representation through kNowledge IntEgration) model, etc.
[0069] In practical application, the processor filters corpus variants, which can not only filter corpus variants generated by the target large model, but also filter corpus at any stage of corpus construction. For example, corpus can be filtered when corpus variants are enhanced and / or the large model is fine-tuned, so as to ensure that the filtered corpus only includes high-quality and task-related corpus, while excluding irrelevant or low-quality output. To filter corpus variants generated by the target large model, a target filtering model needs to be obtained. Specifically, a pre-trained BERT model can be selected as a starting point, which has been trained on a large amount of text and has strong language understanding ability. According to project requirements, different versions of BERT models can be selected, such as BERT-Base, BERT-Large, etc. Although BERT has been pre-trained, in order to improve the filtering quality and relevance, the model can be fine-tuned on relevant corpus of specific tasks to better understand the context of such tasks. Then define positive labels or negative labels for the training corpus used to train the filtering model. The training corpus defined as positive labels is a positive sample, representing high-quality and task-related sentences. The training corpus defined as negative labels is a negative sample, representing low-quality or irrelevant sentences. Collect the labeled training corpus as the training sample of the filtering model, ensure that the training sample covers various contexts and expressions, train the BERT model using the prepared labeled training corpus, and perform binary classification task training to distinguish high-quality / low-quality sentences, and finally obtain the trained target filtering model. The target filtering model can filter the enhanced corpus variants. Specifically, the enhanced corpus variants are filtered through the trained BERT classification model, i.e. the target filtering model. The classification result of each enhanced corpus variant output by the target filtering model is a positive sample (high quality) or a negative sample (low quality or irrelevant). For example, the corpus variant to be enhanced is “turn off ice making”, and the enhanced corpus variant is “<please help me turn off ice making>”, “”, “<stop ice making function>”, “<turn off ice making for me>”, “<stop ice making ice cream>”, and the result of the target filtering model after filtering is “please help me turn off ice making”, “I want to turn off the ice making mode”, “stop ice making function”, “turn off ice making for me”. In addition, the processor can also retrain regularly according to new data and feedback collected during the application of the target filtering model to maintain and improve the filtering quality of the target filtering model.
[0070] In this way, the processor labels the training corpus as positive samples and negative samples, and trains the filtering model using the labeled corpus to obtain a target filtering model. Using a pre-trained model such as BERT as the basis of the filtering model, accurate labeling of high-quality positive samples and low-quality negative samples can provide a clear learning goal for the filtering model, and through supervised learning, the filtering model can learn from the labeled data to accurately distinguish the features of high-quality and low-quality corpus. Therefore, the processor uses the target filtering model trained by the labeled corpus to filter the corpus variants, which can identify and exclude irrelevant or inaccurate corpus, thereby obtaining high-quality optimized corpus.
[0071] Optionally, the processor enhances the corpus variants, comprising: the processor enhances the corpus variants by a simple data enhancement method; wherein the simple data enhancement method comprises one or more of synonym replacement, random insertion, random exchange, random deletion and adding stop words. Synonym replacement refers to selecting key words in a sentence and replacing them with their synonyms, which can enable the model to learn the semantic equivalence between different words, thereby improving its understanding of semantics, especially when dealing with synonyms or near synonyms.
[0072] The EDA (Easy Data Augmentation) technology is used to enhance the generated corpus variants, including one or more operations such as synonym replacement, random insertion, random exchange, random deletion, etc., thereby increasing the diversity of the finally generated corpus variants. Synonym replacement refers to selecting key words in a sentence and replacing them with their synonyms, which can enable the model to learn the semantic equivalence between different words, thereby improving its understanding of semantics, especially when dealing with synonyms or near synonyms. Random insertion refers to randomly inserting words related to the context in the sentence, such as adjectives or adverbs, which can make the sentence more complex and diverse, forcing the model to learn to extract key information from more rich and diverse sentence patterns. Random exchange refers to randomly swapping the order of words in a sentence while maintaining the overall meaning of the sentence, which can enable the model to adapt to and understand various sentence structures and enhance its understanding of sentence meaning under word order changes. Random deletion refers to randomly deleting certain words in a sentence, usually non-key words, which can test and improve the model's ability to recognize the core meaning of a sentence while still correctly understanding the sentence after removing certain information. Adding stop words refers to adding commonly used stop words that usually do not carry the main meaning, such as "of", "is", "and", etc., thereby making the sentence more similar to real-world language habits, which can enable the refrigerator to better understand and process natural language in daily conversations. Specifically, the corpus enhanced by synonym replacement is "Help me start the ice maker mode", the corpus enhanced by random insertion is "Help me open the ice maker mode", the corpus enhanced by random exchange is "Help me open the ice maker mode", the corpus enhanced by random deletion is "Help me open the ice maker mode", and the corpus enhanced by adding stop words is "Please help me open the ice maker mode".
[0073] In this way, the processor enhances the corpus variants by one or more of synonym replacement, random insertion, random exchange, random deletion, and adding stop words, which can increase the diversity of the corpus variants through different transformations, thereby enabling the refrigerator's voice interaction function to accurately understand even when facing unseen expressions and better adapting to natural language changes in the real world, thereby improving the generalization ability of the interaction model.
[0074] Based on the above structure of the refrigerator, as shown in Figure 4 The present disclosure provides a method for constructing a corpus, comprising:
[0075] S01, the processor fine-tunes a large model according to an interaction corpus associated with the function of the intelligent refrigerator, to obtain a target large model.
[0076] S02, the processor inputs the initial corpus into the target large model to generate corpus variants, and filters and / or enhances the corpus variants to obtain optimized corpus.
[0077] S41, the processor sorts the optimized corpus by similarity through a cosine similarity sorting technique, and screens out target corpus meeting preset similarity conditions.
[0078] S42, the processor constructs a corpus based on the target corpus.
[0079] Step S41 is a step after corpus variant enhancement or fine-tuning of a large model, which can screen out the best target corpus from a large number of generated corpus variants for subsequent application. In actual application, corpus vector representation needs to be generated first. Specifically, all generated corpus variants are preprocessed first, including removing punctuation, unifying case, removing stop words, etc., to reduce the influence of noise data on the result; then a pre-trained word embedding model (such as Word2Vec) is used to convert each word in the corpus variant into a vector in a high-dimensional space, and then mathematical methods such as averaging or weighted averaging are used to combine all word vectors in the corpus variant to form a single vector representing the entire corpus variant. After generating the corpus vector representation, the cosine similarity of different corpus variants needs to be calculated. Specifically, for the seed sentence (i.e. the initial corpus) to be enhanced and all generated corpus variants, the single vector representation of each is calculated, and then the cosine similarity formula is used to calculate the similarity between the vector of the initial corpus and the vector of each generated corpus variant. The cosine similarity value ranges from -1 (completely dissimilar) to 1 (completely identical), where a higher value indicates a higher similarity. After calculating the cosine similarity of different corpus variants, the optimal corpus needs to be selected. Specifically, all generated corpus variants are sorted according to the calculated cosine similarity values, where corpus with high similarity is closer to the original semantics of the initial corpus, and corpus with low similarity may introduce too many changes, causing deviation from the original semantics. Finally, corpus variants with a cosine similarity higher than a set similarity threshold are selected as target corpus. The set similarity threshold can be adjusted according to actual application requirements to balance the need for innovation and semantic preservation.
[0080] The method for constructing a corpus provided by the embodiments of the present disclosure, the processor sorts the optimized corpus by similarity through a cosine similarity sorting technique, and screens out target corpus meeting preset similarity conditions. Since the cosine similarity is an effective indicator for measuring the semantic similarity of text, it can accurately reflect the similarity between texts, therefore, by setting a similarity threshold, the proportion of redundant or irrelevant information in the generated corpus variants can be further removed or reduced, thereby controlling the similarity of the screened corpus variants to the initial corpus and ensuring the quality and applicability of the target corpus used to construct the corpus.
[0081] In combination with Figure 5As shown, the embodiment of the present disclosure provides a device 800 for constructing a corpus, comprising a processor 801 and a memory 802. Optionally, the device can further comprise a communication interface 803 and a bus 804. Wherein the processor 801, the communication interface 803, and the memory 802 can complete mutual communication through the bus 804. The communication interface 803 can be used for information transmission. The processor 801 can invoke the logical instructions in the memory 802 to execute the method for constructing a corpus of the above-mentioned embodiments.
[0082] In addition, the logical instructions in the memory 802 described above can be realized in the form of a software functional unit and sold or used as an independent product, which can be stored in a computer readable storage medium.
[0083] The memory 802 as a computer readable storage medium can be used to store software programs, computer executable programs, such as program instructions / modules corresponding to the method in the embodiment of the present disclosure. The processor 801 executes the function application and data processing by running the program instructions / modules stored in the memory 802, that is, realizes the method for constructing a corpus in the above-mentioned embodiments.
[0084] The memory 802 can include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application required by a function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 802 can include a high-speed random access memory, and can also include a non-volatile memory.
[0085] In combination Figure 6 As shown, the embodiment of the present disclosure provides a refrigerator 900, comprising: a refrigerator body, and the above-mentioned device 800 for constructing a corpus. The device 800 for constructing a corpus is installed on the refrigerator body. The installation relationship described herein is not limited to placing in the refrigerator, but also includes installation connection with other components of the refrigerator, including but not limited to physical connection, electrical connection or signal transmission connection, etc. Those skilled in the art can understand that the device 800 for constructing a corpus can be adapted to a feasible refrigerator body, and thus realize other feasible embodiments.
[0086] The embodiment of the present disclosure provides a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions are set to execute the above-mentioned method for constructing a corpus.
[0087] The technical solutions of the embodiments of the present disclosure can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes one or more instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method disclosed in the embodiments of the present disclosure. The aforementioned storage medium can be a non-transitory storage medium, including: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0088] The above description and drawings sufficiently illustrate the embodiments of the present disclosure to enable one skilled in the art to practice them. Other embodiments can include structural, logical, electrical, process, and other changes. The embodiments represent only a few of the possible variations. Individual components and functions are optional unless explicitly required, and the order of operations can be changed. Parts and features of some embodiments can be included or replaced by parts and features of other embodiments. Also, the words used in this application are used only to describe the embodiments and not to limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms as well. Similarly, as used in this application, the term "and / or" refers to any and all possible combinations of one or more associated listed items. In addition, when used in this application, the term "comprise" and its variants "comprises" and / or "comprising" and the like mean the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Without more limitations, the element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, or device that includes the stated element. In this document, each embodiment focuses on the differences from other embodiments, and the same or similar parts between embodiments can be referred to each other. For the method, product, etc. disclosed in the embodiments, if it corresponds to the method part disclosed in the embodiments, the relevant part can be referred to the description of the method part.
[0089] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods for each specific application to realize the described functions, but such implementation should not be considered beyond the scope of the embodiments of the present disclosure. The skilled person can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0090] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the above-described device embodiments are only schematic, for example, the division of the units can only be a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms. The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to implement the embodiments. In addition, each functional unit in the embodiments of the present disclosure can be integrated in one processing unit, or each unit can be a physically independent unit, or two or more units can be integrated in one unit.
[0091] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other processing device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
Claims
1. A method for constructing a corpus, characterized in that, include: The large model is fine-tuned based on the interactive corpus associated with the functions of the smart refrigerator to obtain the target large model; The initial corpus is input into the target large model to generate corpus variants, and the corpus variants are filtered and / or enhanced to obtain optimized corpus; The optimized corpus is sorted by similarity, and the target corpus is selected to construct a corpus.
2. The method according to claim 1, characterized in that, The large model is fine-tuned based on the interaction corpus associated with the functions of the smart refrigerator to obtain the target large model, including: Collect basic interactive corpus related to the functions of smart refrigerators; Data preprocessing is performed on the basic interactive corpus to obtain the first interactive corpus; The large model is fine-tuned based on the first interactive corpus to obtain the target large model.
3. The method according to claim 2, characterized in that, The large model is fine-tuned based on the first interactive corpus to obtain the target large model, including: The first interactive corpus is annotated to obtain the second interactive corpus; Set fine-tuning parameters for the selected large model; The selected large model is fine-tuned using the second interactive corpus according to the set fine-tuning parameters.
4. The method according to claim 1, characterized in that, Input the initial corpus into the target large model to generate corpus variants, including: Input the initial corpus into the target large model to generate initial corpus variants; Input the initial corpus variants into the target large model to generate new corpus variants; If the new corpus variant does not meet the requirements, input the new corpus variant again into the target large model to generate a new initial corpus variant; or, If a new corpus variant meets the requirements, it is identified as a corpus variant.
5. The method according to claim 1, characterized in that, The corpus variants were filtered, including: The training corpus is labeled as positive and negative samples; The filtering model is trained using the labeled corpus to obtain the target filtering model; The optimized corpus is obtained by filtering the variants of the corpus using a target filtering model.
6. The method according to claim 1, characterized in that, Enhancements to the corpus variants include: Enhance the corpus variants using simple data augmentation methods; Simple data augmentation methods include one or more of the following: synonym replacement, random insertion, random swapping, random deletion, and adding stop words.
7. The method according to any one of claims 1 to 6, characterized in that, The optimized corpus is sorted by similarity, and the target corpus is selected to construct a corpus, including: The optimized corpus is sorted by similarity using cosine similarity ranking technology, and target corpus that meets the preset similarity conditions is selected. A corpus is constructed based on the target corpus.
8. An apparatus for constructing a corpus, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to perform the method of constructing a corpus as described in any one of claims 1 to 7 when executing the program instructions.
9. A refrigerator, characterized in that, include: Refrigerator body; The apparatus for constructing a corpus as described in claim 8 is installed on the refrigerator body.
10. A computer-readable storage medium storing program instructions, characterized in that, When the program instructions are executed, they cause the computer to perform the method for constructing a corpus as described in any one of claims 1 to 7.
Citation Information
Cited By
Target operating system-oriented intention analysis and intelligent control method and system
CN122222038A
A method and system for intent parsing and intelligent control of a target operating system
CN122222038B