Cooking video course demonstration method and system, electronic equipment and storage medium

By introducing a quantitative parameter model in the cooking video tutorial, converting fuzzy parameters into quantitative parameters, and combining voice broadcasting and user manipulation instructions, the problems of strong steps continuity, high language description and interactive channel conflict in the existing cooking video tutorial are solved, achieving a more efficient and more accurate cooking learning experience.

CN120223977APending Publication Date: 2025-06-27广州千阅传媒有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510351726.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing cooking video tutorials have problems such as strong step continuity, high language description subjectivity, and conflict in interaction channels, resulting in low learning efficiency, heavy memory burden and high risk of dish pollution.

Method used

By obtaining the cooking video tutorial that has been segmented according to the operation steps, extracting the voice for recognition and obtaining the recipe text, inputting the quantitative parameter model to obtain the quantization parameters of the fuzzy parameters, and superimposing it into the video tutorial to realize the playback control of voice broadcast and user manipulation instructions.

Benefits of technology

It improves users' precise time, time, temperature and quantity control during cooking, reduces learning difficulty and risk of dish pollution, and improves the learning efficiency of cooking video tutorials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223977A_ABST
    Figure CN120223977A_ABST
Patent Text Reader

Abstract

The invention relates to a cooking video course demonstration method and system, electronic equipment and a storage medium. The method comprises the following steps: acquiring a cooking video course segmented according to operation steps; extracting voice of the cooking video course, and identifying the voice to obtain a corresponding menu text; inputting the menu text into a pre-trained and constructed quantization parameter model to obtain quantization parameters corresponding to the fuzzy parameters in the menu text; superposing the quantization parameter to a corresponding segment of a cooking video course to obtain a new cooking video course; and obtaining a control instruction of the user, and performing playing control on the new cooking video course according to the control instruction of the user. The fuzzy parameters are converted into the quantized parameters, and when the new cooking video course is played to the corresponding fragment, the corresponding quantized parameters are synchronously displayed, so that a user can more easily grasp cooking skills, and accurate time control and temperature control are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and in particular, to a method, system, electronic device, and storage medium for demonstrating a cooking video tutorial. Background Art

[0002] In recent years, the rapid development of fresh food e-commerce, community group buying, etc. has significantly lowered the threshold of home cooking, and the convenience of obtaining ingredients has been greatly improved. At the same time, with the growth of consumers' demand for dietary health, home cooking has gradually shifted towards the pursuit of dish diversity, nutritional collocation, and refinement of production processes. The popularization of intelligent kitchen utensils such as ovens, air fryers, and blenders has further broadened the types of dishes that can be tried at home (such as baking, molecular cuisine, etc.), but it has also made the cooking process more complex, and users need to master skills such as multi-step operations, precise control of time, temperature, and quantity, and the use of complex tools.

[0003] There are mainly two ways to learn new dish cooking methods. One is cooking video tutorials, and the other is illustrated recipes.

[0004] Currently, the following problems exist in the way of cooking video tutorials:

[0005] 1. Excessive step continuity: The video is not segmented according to operation nodes, and users need to memorize the whole process or frequently operate pause / rewind during cooking (≥5 pauses per dish on average), resulting in low learning efficiency and heavy memory burden;

[0006] 2. Strong subjectivity in language description in recipe videos: relying on vague terms (such as "appropriate amount of salt", "medium-low heat"), lacking quantitative definitions, and it is difficult for users to precisely control time, temperature, and quantity;

[0007] 3. Interaction channel conflict: During cooking, users' hands touch ingredients or kitchen utensils, and then frequently operate electronic devices, resulting in a high risk of dish contamination. Summary of the Invention

[0008] In order to solve or partially solve the above technical problems, the present invention provides a method, system, electronic device, and storage medium for demonstrating a cooking video tutorial. By using a model, the quantitative parameters corresponding to the fuzzy parameters in the recipe text corresponding to the cooking video tutorial are obtained, and the quantitative parameters are superimposed on the corresponding segments of the cooking video tutorial, so that users can easily precisely control time, temperature, and quantity.

[0009] In a first aspect, an embodiment of the present invention provides a method for demonstrating a cooking video tutorial. The method includes:

[0010] Obtain a cooking video tutorial that has been segmented according to operation steps;

[0011] Extract the voice of the cooking video tutorial, and perform speech recognition on the voice to obtain the corresponding recipe text;

[0012] Input the recipe text into a pre-trained quantization parameter model to obtain the quantization parameters corresponding to the fuzzy parameters in the recipe text;

[0013] Overlay the quantization parameters on the corresponding segments of the cooking video tutorial to obtain a new cooking video tutorial;

[0014] Obtain the user's control instruction and perform playback control on the new cooking video tutorial according to the user's control instruction.

[0015] In one implementation, the quantization parameter model is a large language model.

[0016] In one implementation, the quantization parameters output by the quantization parameter model adopt a structured JSON format.

[0017] In one implementation, the control instruction includes a voice control instruction; obtaining the user's control instruction includes:

[0018] Obtain the user's voice, and perform voice recognition on the user's voice to obtain a voice control instruction.

[0019] In one implementation, the method further includes:

[0020] When the new cooking video tutorial is played to the frame where the quantization parameters are displayed, the quantization parameters are synchronously announced by voice.

[0021] In a second aspect, an embodiment of the present invention provides a cooking video tutorial demonstration system. The system includes:

[0022] A video acquisition module for acquiring a cooking video tutorial that has been segmented according to operation steps;

[0023] A recipe text recognition module for extracting the voice of the cooking video tutorial and performing voice recognition on the voice to obtain the corresponding recipe text;

[0024] A quantization parameter module for inputting the recipe text into a pre-trained quantization parameter model to obtain the quantization parameters corresponding to the fuzzy parameters in the recipe text;

[0025] A new video generation module for overlaying the quantization parameters on the corresponding segments of the cooking video tutorial to obtain a new cooking video tutorial;

[0026] A playback control module for obtaining the user's control instruction and performing playback control on the new cooking video tutorial according to the user's control instruction.

[0027] In one implementation, the quantization parameter model is a large language model.

[0028] In one embodiment, the quantization parameter output by the quantization parameter model adopts a structured JSON format.

[0029] In one embodiment, the control instruction includes a voice control instruction; the playback control module is further configured to obtain the user's voice and identify the voice control instruction from the user's voice.

[0030] In one embodiment, the system further includes:

[0031] A voice broadcast module, configured to synchronously perform voice broadcast on the quantization parameter when the new cooking video tutorial is played to the frame displaying the quantization parameter.

[0032] In a third aspect, the present invention provides an electronic device, which includes:

[0033] At least one processor; and

[0034] A memory communicatively connected to the at least one processor; wherein,

[0035] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the above-mentioned cooking video tutorial demonstration method.

[0036] In a fourth aspect, the present invention provides a computer-readable storage medium, which stores computer instructions for causing a processor to implement the above-mentioned cooking video tutorial demonstration method when executed.

[0037] In this embodiment, by obtaining a cooking video tutorial segmented according to operation steps; extracting the voice of the cooking video tutorial and identifying the corresponding recipe text from the voice; inputting the recipe text into a pre-trained and constructed quantization parameter model to obtain the quantization parameter corresponding to the fuzzy parameter in the recipe text; superimposing the quantization parameter on the corresponding segment of the cooking video tutorial to obtain a new cooking video tutorial; obtaining the user's control instruction and performing playback control on the new cooking video tutorial according to the user's control instruction. The conversion of the fuzzy parameter into a quantization parameter is realized, and when the new cooking video tutorial is played to the corresponding segment, the corresponding quantization parameter will be synchronously displayed, making it easier for the user to master cooking skills and achieve precise timing, temperature control, and quantity control. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0039] Figure 1 It is a flowchart of a cooking video tutorial demonstration method provided in the first embodiment of the present invention;

[0040] Figure 2 It is a schematic structural diagram of a cooking video tutorial demonstration system provided in the second embodiment of the present invention;

[0041] Figure 3 It is a schematic structural diagram of an electronic device provided in the embodiment of the present invention. Detailed implementation manners

[0042] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0043] Figure 1 It is a flowchart of a cooking video tutorial demonstration method provided in the first embodiment of the present invention. This embodiment can be applied to a cooking video tutorial demonstration system, such as Figure 1 As shown, this embodiment may include the following steps:

[0044] Step 101, obtain a cooking video tutorial segmented according to operation steps.

[0045] In this step, the cooking video tutorial demonstration system can provide a configuration interface for users or administrators to upload a cooking video tutorial segmented according to operation steps. The cooking video tutorial demonstration system can also provide a management background, where the administrator uploads the cooking video tutorial to the management background and segments the cooking video tutorial according to operation steps in the management background. For example, it is segmented according to operation steps such as ingredient preparation, marinating, frying 1, frying 2, etc., so as to obtain a cooking video tutorial segmented according to operation steps.

[0046] Step 102, extract the voice of the cooking video tutorial and perform voice recognition to obtain the corresponding recipe text.

[0047] In this step, the voice (audio stream) of the cooking video tutorial can be extracted through the ffmpeg tool. After obtaining the voice, voice recognition is performed on the voice to obtain the corresponding recipe text.

[0048] Step 103, input the recipe text into a pre-trained and constructed quantization parameter model to obtain the quantization parameters corresponding to the fuzzy parameters in the recipe text.

[0049] In this step, the quantization parameters corresponding to the fuzzy parameters in the recipe text of the cooking video tutorial are obtained through the quantization parameter model. For example, if the recipe text states that an appropriate amount of salt is added during frying method 1, the quantization parameter model can obtain the quantization parameter "usage: 5 grams" corresponding to this fuzzy parameter "appropriate amount", that is, the usage of salt added during frying method 1 is 5 grams.

[0050] The training and construction of the quantization parameter model may include the following steps:

[0051] 1. Data collection and annotation

[0052] Collect a large number of recipe texts from various sources, including cooking websites (such as AllRecipes, Xiachufang), cooking books, social media, etc., to ensure that the data contains various types of fuzzy parameter expressions.

[0053] Meanwhile, collect some recipe texts with already annotated quantization parameters, and / or, invite chefs or cooking experts to annotate the quantization parameters corresponding to the fuzzy parameters (for example, annotate "one spoonful of oil" as "15 ml" and "one spoonful of sugar" as "20 grams") to obtain the recipe texts as the basis for training.

[0054] Annotation specifications

[0055] Term classification: Classify according to cooking operations (such as cooking temperature, usage, time, state, etc.).

[0056] Parameter format: Numerical range (grams), time interval (seconds / minutes).

[0057] Context association: Record the context in which the term appears. Considering the context association can more accurately determine the quantization parameters corresponding to the fuzzy parameters.

[0058] 2. Data preprocessing

[0059] Text cleaning:

[0060] Remove irrelevant symbols, such as advertisements, emojis, special characters, etc.

[0061] Unify units, such as "one spoonful = 15 ml", "one cup = 240 ml".

[0062] Structured storage:

[0063] For example

[0064] {

[0065] "text": "After the oil is hot, turn to medium-low heat and add an appropriate amount of salt",

[0066] "labels": {

[0067] "operation": "Heat adjustment",

[0068] "term": "Medium - low heat",

[0069] "value": "180℃ - 200℃",

[0070] "context": "Stir - frying"

[0071] }

[0072] }

[0073] 3. Model Selection and Fine - Tuning

[0074] Base model: Large language models that support long - text generation such as GPT - 4, Llama - 2, etc. can be selected.

[0075] Domain adaptation: Continue pre - training on cooking - domain corpora (such as recipes, ingredient encyclopedias) to enhance term understanding.

[0076] Fine - tuning strategy

[0077] Task design: Define the problem as term - to - parameter mapping. The input is recipe text, and the output is structured parameters.

[0078] Prompt template:

[0079] Input: {Recipe text}

[0080] Task: Extract the following information:

[0081] a. Vague parameters (such as "appropriate amount", "one spoon").

[0082] b. Corresponding quantified parameters (time, dosage, etc.).

[0083] c. Context - related scenarios (such as stir - frying, stewing).

[0084] 4. Model Training and Optimization

[0085] Training steps:

[0086] Domain - adaptation pre - training: Continue training the model with cooking corpora to learn professional terms (such as "blanching", "reducing the sauce").

[0087] Supervised fine - tuning (SFT): Train on the labeled dataset. The input is recipe text, and the output is parameter key - value pairs.

[0088] Reinforcement learning (RLHF): Introduce expert feedback to optimize the rationality of generating quantified parameters (for example, "high heat" should not exceed the smoke point temperature).

[0089] 5. Model Evaluation and Validation

[0090] Evaluation Metrics

[0091] Accuracy: The matching accuracy of fuzzy parameters to quantization parameters (e.g., whether "a spoonful → 15 grams" is correct).

[0092] 6. Structured Output of Quantization Parameters

[0093] For example

[0094] {

[0095] "terms":

[0096] {

[0097] "term": "appropriate amount of salt",

[0098] "type": "usage amount",

[0099] "value": "5 grams",

[0100] "context": "500 grams of ingredients"

[0101] }

[0103] }

[0104] Step 104: Overlay the quantization parameters onto the corresponding segments of the cooking video tutorial to obtain a new cooking video tutorial.

[0105] In this step, overlay the obtained quantization parameters onto the corresponding segments of the cooking video tutorial. When the new cooking video tutorial plays to the corresponding segments, the corresponding quantization parameters will be displayed synchronously, enabling users to more easily grasp cooking techniques and achieve precise control of time, temperature, and quantity.

[0106] Step 105: Obtain the user's manipulation instructions and control the playback of the new cooking video tutorial according to the user's manipulation instructions.

[0107] In this step, the cooking video tutorial demonstration system controls the playback of the new cooking video tutorial according to the user's manipulation instructions, such as replaying the previous segment, pausing, continuing, replaying, starting to play, etc.

[0108] ​To avoid the operation of the user's hand from contaminating the dishes in the cooking video tutorial demonstration system, in this embodiment, the control instruction includes a voice control instruction. After the cooking video tutorial demonstration system obtains the user's voice, it recognizes the user's voice to obtain a voice control instruction, and then controls the playback of the new cooking video tutorial according to the user's voice control instruction. The user can interact with the system only through voice (such as "next step", "repeat the current step") to achieve precise manual-free control of the new cooking teaching video.

[0109] The cooking video tutorial demonstration system can use a directional microphone array to receive the user's voice in real time and access the system through a voice input interface. The directional microphone array is a hardware that realizes directional sound pickup and noise suppression by the collaborative work of multiple microphone units and utilizes the spatial characteristics of sound waves. It processes the differences in sound wave signals received by multiple microphones through algorithms to form a "sound beam" in a specific direction, enhancing the voice in the target direction while suppressing interference noise in other directions. Thus, it can focus on the user's voice and suppress background noise, such as the range hood, frying sound, etc.

[0110] The user's voice can be recognized using cloud voice recognition services. Specifically, the user's voice is sent to the cloud voice recognition service for voice recognition, keywords are matched, and it is converted into a recognizable voice control instruction.

[0111] In one implementation, the method further includes:

[0112] Step 105a, when the new cooking video tutorial plays to the frame showing the quantization parameters, the quantization parameters are synchronously announced in voice.

[0113] In this embodiment, by obtaining a cooking video tutorial segmented according to the operation steps; extracting the voice of the cooking video tutorial, recognizing the voice to obtain the corresponding recipe text; inputting the recipe text into a pre-trained and constructed quantization parameter model to obtain the quantization parameters corresponding to the fuzzy parameters in the recipe text; superimposing the quantization parameters on the corresponding segments of the cooking video tutorial to obtain a new cooking video tutorial; obtaining the user's control instruction, and controlling the playback of the new cooking video tutorial according to the user's control instruction. It realizes the conversion of fuzzy parameters into quantization parameters. When the new cooking video tutorial plays to the corresponding segment, the corresponding quantization parameters will be synchronously displayed, making it easier for the user to master cooking skills and achieve precise control of time, temperature, and quantity.

[0114] Corresponding to the cooking video tutorial demonstration method in the present invention, the present invention also provides a cooking video tutorial demonstration system, Figure 2 which is a structural schematic diagram of a cooking video tutorial demonstration system. As Figure 2 shown, the cooking video tutorial demonstration system includes:

[0115] The video acquisition module 201 is configured to acquire a segmented cooking video tutorial according to the operation steps;

[0116] The recipe text recognition module 202 is configured to extract the voice of the cooking video tutorial and recognize the voice to obtain the corresponding recipe text;

[0117] The quantization parameter module 203 is configured to input the recipe text into a pre-trained and constructed quantization parameter model to obtain the quantization parameters corresponding to the fuzzy parameters in the recipe text;

[0118] The new video generation module 204 is configured to superimpose the quantization parameters on the corresponding segments of the cooking video tutorial to obtain a new cooking video tutorial;

[0119] The playback control module 205 is configured to acquire a user's manipulation instruction and perform playback control on the new cooking video tutorial according to the user's manipulation instruction.

[0120] In one embodiment, the quantization parameter model is a large language model.

[0121] In one embodiment, the quantization parameters output by the quantization parameter model adopt a structured JSON format.

[0122] In one embodiment, the manipulation instruction includes a voice manipulation instruction; the playback control module 205 is further configured to acquire a user's voice and recognize the user's voice to obtain a voice manipulation instruction.

[0123] In one embodiment, the system further includes:

[0124] The voice broadcast module is configured to synchronously perform voice broadcast on the quantization parameters when the new cooking video tutorial is played to a frame displaying the quantization parameters.

[0125] A cooking video tutorial demonstration system provided by an embodiment of the present invention can execute a cooking video tutorial demonstration method provided by any embodiment of the present invention, and has function modules and beneficial effects corresponding to the execution of the method.

[0126] Figure 3 FIG. shows a schematic structural diagram of an electronic device 30 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described herein and / or claimed.

[0127] As shown Figure 3 in FIG. 1, the electronic device 30 includes at least one processor 31 and a memory communicatively connected to the at least one processor 31, such as a read-only memory (ROM) 32, a random access memory (RAM) 33, etc. The memory stores a computer program executable by the at least one processor. The processor 31 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 32 or the computer program loaded from the storage unit 38 into the random access memory (RAM) 33. In the RAM 33, various programs and data required for the operation of the electronic device 30 can also be stored. The processor 31, the ROM 32, and the RAM 33 are connected to each other via a bus 34. An input / output (I / O) interface 35 is also connected to the bus 34.

[0128] A plurality of components in the electronic device 30 are connected to the I / O interface 35, including: an input unit 36, such as a keyboard, a mouse, etc.; an output unit 37, such as various types of displays, speakers, etc.; a storage unit 38, such as a magnetic disk, an optical disc, etc.; and a communication unit 39, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 39 allows the electronic device 30 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0129] The processor 31 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 31 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 31 executes the various methods and processes described above, such as the cooking video tutorial demonstration method.

[0130] In some embodiments, the cooking video tutorial demonstration method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 38. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 30 via the ROM 32 and / or the communication unit 39. When the computer program is loaded into the RAM 33 and executed by the processor 31, one or more steps of the cooking video tutorial demonstration method described above can be executed. Alternatively, in other embodiments, the processor 31 can be configured to execute the cooking video tutorial demonstration method by any other appropriate means (e.g., by means of firmware).

[0131] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0132] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0133] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0134] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0135] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0136] The computing system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0137] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and this is not limited herein.

[0138] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A cooking video tutorial demonstration method, characterized in that: The method comprises: Get cooking video tutorials segmented into steps; Extract the speech from the cooking video tutorial and recognize the speech to get the corresponding recipe text; Input the recipe text into the pre-trained quantitative parameter model to obtain the quantitative parameters corresponding to the fuzzy parameters in the recipe text; superimposing the quantized parameters onto corresponding segments of the cooking video tutorial to obtain a new cooking video tutorial; Obtain the user's control instructions and control the playback of the new cooking video tutorial according to the user's control instructions.

2. The method according to claim 1, characterized in that: The quantization parameter model is a large language model.

3. The method according to claim 2, characterized in that: The quantization parameters output by the quantization parameter model are in a structured JSON format.

4. The method according to claim 1, characterized in that: The control instruction includes a voice control instruction; and the obtaining of the user's control instruction includes: Acquire user voice, recognize user voice and get voice control instructions.

5. The method according to claim 1, characterized in that Also includes: When the new cooking video tutorial is played to the frame displaying the quantization parameter, the quantization parameter is simultaneously voice broadcasted.

6. A cooking video tutorial demonstration system, characterized in that: The system comprises: A video acquisition module is used to acquire cooking video tutorials that have been segmented according to operation steps; The recipe text recognition module is used to extract the voice of the cooking video tutorial and recognize the voice to obtain the corresponding recipe text; A quantization parameter module is used to input the recipe text into a pre-trained quantization parameter model to obtain the quantization parameters corresponding to the fuzzy parameters in the recipe text; A new video generation module, used for superimposing the quantization parameters onto corresponding segments of the cooking video tutorial to obtain a new cooking video tutorial; The playback control module is used to obtain the user's control instructions and control the playback of the new cooking video tutorial according to the user's control instructions.

7. The system according to claim 6, characterized in that: The control instructions include voice control instructions; the playback control module is also used to obtain user voice and recognize the user voice to obtain the voice control instructions.

8. The system according to claim 6, characterized in that The system further comprises: The voice broadcast module is used to synchronously voice broadcast the quantization parameters when the new cooking video tutorial is played to the frame where the quantization parameters are displayed.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the cooking video tutorial demonstration method described in any one of claims 1-5.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the cooking video tutorial demonstration method described in any one of claims 1-5 when executed.