Home metadata construction method and device and control page generation method and device
Through speech recognition and large language model analysis of user instructions, home metadata is constructed, which solves the problem of time-consuming, labor-intensive and error-based entry in smart home systems, and realizes efficient and accurate home metadata management and device control.
Patent Information
- Application Number
- CN202510481601.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-12
AI Technical Summary
In existing smart home systems, the entry of home metadata is time-consuming and labor-intensive and error-prone, and the speech recognition accuracy is low, making it impossible to understand the user's complex intentions.
Through voice recognition technology, user instructions are converted into text, and keywords are parsed using large language models, match cloud databases to build home metadata, and generate control pages.
It improves the efficiency of home metadata generation, reduces manual entry errors, and realizes accurate identification and device control of user fuzzy intentions.
Smart Images

Figure CN120472894A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communications, and in particular, to a method and device for constructing family metadata, and a method and device for generating a control page. Background Art
[0002] Throughout the development of smart home systems and related IoT technologies, the management and maintenance of home metadata has been crucial for improving user experience and technical efficiency. Home metadata includes detailed information about physical spaces (e.g., rooms), IoT devices (e.g., smart light bulbs, air conditioners, security cameras), and the relationships between these devices and spaces.
[0003] However, in the related art, existing smart home systems usually rely on users to manually enter metadata such as home layout, room names, and device information. Users need to enter the name of each room and detailed information of all devices in the room one by one, including device ID, type, brand, etc. This is not only time-consuming, but may also lead to a decline in user experience due to the cumbersome entry process. During the manual entry process, users may enter incorrect data due to spelling errors, inaccurate memory of device information, etc., which in turn affects the control of the device and the normal operation of the system. In addition, different users may use different names for the same device, such as "desk lamp", "bedside lamp" or "study lamp". This inconsistency makes it difficult for the system to uniformly manage and identify devices.
[0004] Furthermore, in related technologies, single speech recognition technologies may not accurately understand the user's intent, significantly reducing speech recognition accuracy. Even if speech recognition technology can convert speech into text, existing technologies may not be able to interpret ambiguous or complex user control intentions when parsing text commands due to a lack of advanced natural language processing capabilities.
[0005] Regarding the related technologies, manually inputting family metadata is not only time-consuming and labor-intensive, but also prone to errors in family metadata. No effective solution has yet been proposed. Summary of the Invention
[0006] The embodiments of the present application provide a method and device for constructing family metadata, and a method and device for generating a control page, so as to at least solve the problem in the related art that manual input of family metadata is not only time-consuming and labor-intensive, but also prone to errors in family metadata.
[0007] According to one embodiment of the present application, a method for constructing family metadata is provided, which is applied to a terminal device, including: obtaining a voice command issued by a target object, and converting the voice command into a first text through voice recognition technology, or obtaining a second text sent by the target object; parsing the first text or the second text through a large language model to obtain multiple first keywords in the first text or the second text, wherein the multiple first keywords correspond one-to-one to multiple spaces in the target area where the target object is located or multiple Internet of Things devices in the multiple spaces; when the multiple first keywords successfully match the target keywords in the database of the cloud server, constructing family metadata based on the multiple first keywords, wherein the family metadata is at least used to characterize the positional relationship between the multiple Internet of Things devices and the multiple spaces.
[0008] In an exemplary embodiment, after constructing family metadata based on the multiple first keywords, the method also includes: determining multiple second keywords contained in the multiple first keywords corresponding to the multiple spaces, and multiple third keywords corresponding to the multiple IoT devices in the multiple spaces; generating multiple first controls corresponding to the multiple second keywords and multiple second controls corresponding to the multiple third keywords through front-end page generation rules; determining the second position relationship between the multiple first controls and the multiple second controls based on the first position relationship between the multiple spaces and the multiple IoT devices; and generating a first control page of the terminal device based on the multiple first controls, the multiple second controls and the second position relationship.
[0009] In an exemplary embodiment, after generating the first control page of the terminal device according to the multiple first controls, the multiple second controls and the second position relationship, the method further includes: receiving an adjustment instruction sent by the target object, and obtaining a fourth keyword corresponding to the adjustment instruction through the Internet of Things interface, wherein the adjustment instruction is used to add or delete the fourth keyword in the family metadata; according to the adjustment instruction, adjusting the control corresponding to the fourth keyword on the first control page through the control interface of the front-end page, wherein the front-end page is used to display the first control page.
[0010] In an exemplary embodiment, before parsing the first text or the second text through a large language model and obtaining multiple first keywords in the first text or the second text, the method also includes: changing the format of the third text to obtain a third text in a standard format, wherein the third text is the first text or the second text; performing noise filtering on the third text in the standard format to obtain a clarified third text; and extracting multiple fifth keywords from the clarified third text, wherein the multiple fifth keywords include at least one of the following: the multiple first keywords, a sixth keyword, and the sixth keyword is used to indicate the time or action corresponding to the third text.
[0011] In an exemplary embodiment, after constructing family metadata based on the multiple first keywords, the method further includes: obtaining a scene control instruction issued by the target object from the first text or the second text, wherein the scene control instruction is used to control the target Internet of Things device to perform a target action; determining a target space from the multiple spaces based on the scene control instruction, and determining a second control page corresponding to the target space; and displaying the control instruction on the second control page.
[0012] In an exemplary embodiment, after constructing family metadata based on the multiple first keywords, the method also includes: obtaining the status of multiple IoT devices in the multiple spaces through the IoT interface; when there is a first IoT device among the multiple IoT devices and the status of the first IoT device is stopped, deleting the third keyword corresponding to the first IoT device from the family metadata; when a second IoT device is added to the multiple spaces, updating the family metadata according to the seventh keyword corresponding to the second IoT device.
[0013] In an exemplary embodiment, after parsing the first text or the second text through a large language model and obtaining multiple first keywords in the first text or the second text, the method further includes: when there is an eighth keyword among the multiple first keywords that fails to match multiple keywords in the database of the cloud server, receiving an error message sent by the cloud server, wherein the error message includes the reason why the eighth keyword fails to match the multiple keywords.
[0014] According to another embodiment of the embodiment of the present application, a device for constructing family metadata is also provided, including: an acquisition module, used to obtain voice instructions issued by a target object, and convert the voice instructions into a first text through voice recognition technology, or obtain a second text sent by the target object; a parsing module, used to parse the first text or the second text through a large language model, and obtain multiple first keywords in the first text or the second text, wherein the multiple first keywords correspond one-to-one to multiple spaces in the target area where the target object is located or multiple Internet of Things devices in the multiple spaces; a construction module, used to construct family metadata based on the multiple first keywords when the multiple first keywords successfully match the target keywords in the database of the cloud server, wherein the family metadata is at least used to characterize the positional relationship between the multiple Internet of Things devices and the multiple spaces.
[0015] According to one embodiment of the present application, a method for generating a control page is provided, which is applied to a terminal device, including: obtaining a voice instruction issued by a target object, and converting the voice instruction into a first text through voice recognition technology, or obtaining a second text sent by the target object; parsing the first text or the second text through a large language model to obtain multiple first keywords in the first text or the second text, wherein the multiple first keywords correspond one-to-one to multiple spaces in the target area where the target object is located or multiple Internet of Things devices in the multiple spaces; when the multiple first keywords successfully match the target keywords in the database of the cloud server, constructing family metadata based on the multiple first keywords, wherein the family metadata is at least used to characterize the positional relationship between the multiple Internet of Things devices and the multiple spaces; generating a target control corresponding to the family metadata through a front-end page generation rule, and generating a first control page based on the target control.
[0016] According to another embodiment of the embodiment of the present application, a control page generation device is also provided, including: an acquisition module, used to acquire voice instructions issued by a target object, and convert the voice instructions into a first text through voice recognition technology, or acquire a second text sent by the target object; a parsing module, used to parse the first text or the second text through a large language model, and acquire multiple first keywords in the first text or the second text, wherein the multiple first keywords correspond one-to-one to multiple spaces in the target area where the target object is located or multiple Internet of Things devices in the multiple spaces; a construction module, used to construct family metadata based on the multiple first keywords when the multiple first keywords successfully match the target keywords in the database of the cloud server, wherein the family metadata is at least used to characterize the positional relationship between the multiple Internet of Things devices and the multiple spaces; a generation module, used to generate a target control corresponding to the family metadata through a front-end page generation rule, and generate a first control page based on the target control.
[0017] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the above method when running.
[0018] According to another aspect of an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the method through the computer program.
[0019] In an embodiment of the present application, a voice command issued by a target object is obtained, and the voice command is converted into a first text through voice recognition technology, or a second text sent by the target object is obtained; the first text or the second text is parsed through a large language model to obtain multiple first keywords in the first text or the second text, wherein the multiple first keywords correspond one-to-one to multiple spaces in the target area where the target object is located or multiple IoT devices in multiple spaces; when the multiple first keywords are successfully matched with the target keywords in the database of the cloud server, family metadata is constructed based on the multiple first keywords, wherein the family metadata is at least used to characterize the positional relationship between the multiple IoT devices and the multiple spaces. By adopting the above scheme, the multiple first keywords in the voice command issued by the target object are accurately matched with the target keywords in the database through a large language model, and family metadata is generated based on the multiple first keywords, thereby solving the problem in the related art that manually inputting family metadata is not only time-consuming and labor-intensive, but also prone to errors in family metadata, effectively improving the generation efficiency of family metadata, and realizing accurate recognition of the fuzzy intentions of the target object. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0021] Figure 1 This is a hardware structure block diagram of a terminal device for a method for constructing family metadata according to an embodiment of the present application;
[0022] Figure 2 is a flowchart of a method for constructing family metadata according to an embodiment of the present application;
[0023] Figure 3 This is a structural block diagram of a system for generating a control page according to an embodiment of the present application;
[0024] Figure 4 This is a first schematic diagram of a method for constructing family metadata according to an embodiment of the present application;
[0025] Figure 5 is a second schematic diagram of a method for constructing family metadata according to an embodiment of the present application;
[0026] Figure 6 This is a structural block diagram of a device for constructing family metadata according to an embodiment of the present application;
[0027] Figure 7 is a flow chart of a method for generating a control page according to an embodiment of the present application;
[0028] Figure 8 This is a structural block diagram of a device for generating a control page according to an embodiment of the present application. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0031] The method embodiments provided in the embodiments of the present application can be executed in a terminal device or a similar computing device. Taking running on a terminal device as an example, Figure 1 This is a hardware structure block diagram of a terminal device for a method of constructing family metadata in an embodiment of the present application. Figure 1 As shown, the terminal device may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data. In an exemplary embodiment, the terminal device may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above terminal device. Figure 1 More or fewer components than shown, or with Figure 1 Equivalent functions or comparisons shown Figure 1 Shown are different configurations with more functionality.
[0032] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the terminal device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0033] The transmission device 106 is used to receive or send data via a network. A specific example of the aforementioned network may include a wireless network provided by a communications provider of the terminal device. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0034] In this embodiment, a method for constructing family metadata is provided, which is applied to a terminal device. Figure 2 : is a flowchart of a method for constructing family metadata according to an embodiment of the present application, the process including the following steps:
[0035] Step S202: obtaining a voice command sent by a target object, and converting the voice command into a first text using a voice recognition technology, or obtaining a second text sent by the target object;
[0036] Step S204: parsing the first text or the second text using a large language model to obtain a plurality of first keywords in the first text or the second text, wherein the plurality of first keywords correspond one-to-one to a plurality of spaces in a target area where the target object is located or a plurality of IoT devices in the plurality of spaces;
[0037] Step S206: When the multiple first keywords successfully match the target keywords in the database of the cloud server, construct family metadata based on the multiple first keywords, wherein the family metadata is at least used to characterize the positional relationship between the multiple IoT devices and the multiple spaces.
[0038] It's important to note that the aforementioned cloud servers store user input uploaded by all devices. This information is processed and converted into structured data, including but not limited to user preferences, historical control commands, home layout information, device lists and their status. This centralized data management facilitates the unified maintenance and updating of home metadata, ensuring the consistency and accuracy of all device and space information.
[0039] When the user sends a control command by voice or text, the terminal device first performs preliminary voice recognition or text analysis, and then uploads the extracted key information (i.e., multiple first keywords) to the cloud server. The large language model in the cloud server further deeply analyzes these keywords to understand the user's true intentions and request details. For example, the user says "There is a ceiling lamp, two grille lights, and two embedded spotlights in the living room." The large language model will parse out the room named "living room" and the information of 5 specific lighting products in the room. The large language model is used to match the parsed results with the target keywords in the database. If the match is successful, the above-mentioned multiple first metadata are stored in the database of the cloud server, and family metadata is constructed based on the multiple first keywords.
[0040] Through the above steps, a voice command issued by the target object is obtained, and the voice command is converted into a first text through voice recognition technology, or a second text sent by the target object is obtained; the first text or the second text is parsed using a large language model to obtain multiple first keywords in the first text or the second text, wherein the multiple first keywords correspond one-to-one with multiple spaces in the target area where the target object is located or multiple IoT devices in the multiple spaces; when the multiple first keywords successfully match the target keywords in the database of the cloud server, family metadata is constructed based on the multiple first keywords, wherein the family metadata is used to at least characterize the positional relationship between the multiple IoT devices and the multiple spaces. The above method solves the problem in the related art of manually inputting family metadata, which is not only time-consuming and labor-intensive, but also prone to errors in the family metadata.
[0041] Optionally, suppose the user says to the terminal device: My house has two bedrooms and a living room, a master bedroom, a second bedroom, a living room, and a bathroom. The living room has a ceiling lamp, two grille lamps, and two embedded spotlights. The second bedroom has four downlights. The master bedroom has a ceiling lamp and four spotlights. The bathroom has a ceiling lamp. After receiving the user's voice command, the terminal device converts it into a first text, and parses the first text through a large language model to obtain multiple first keywords in the first text, for example, master bedroom, second bedroom, living room, bathroom, ceiling lamp, grille lamp, spotlight, downlight. In the parsing process, the large language model not only identifies multiple first keywords, but also understands the relationship between the above multiple first keywords (such as which devices belong to which room and how many devices are in each room), and extracts this information through a dedicated prompt or instruction set.
[0042] The multiple first keywords are matched against the target keywords in the database. A successful match indicates that the multiple first keywords exist in the database. The target keywords are a list of device and space names pre-set in the database. At this point, household metadata is constructed based on the multiple first keywords. For example, in the living room, there are: ceiling light {quantity: 1}, grille light {quantity: 2}, spotlight {quantity: 2}; in the second bedroom, there are downlights {quantity: 4}; in the master bedroom, there are ceiling light {quantity: 1}, spotlights {quantity: 4}; and in the bathroom, there are ceiling light {quantity: 1}.
[0043] In an exemplary embodiment, after constructing family metadata based on the multiple first keywords, the method also includes: determining multiple second keywords contained in the multiple first keywords corresponding to the multiple spaces, and multiple third keywords corresponding to the multiple IoT devices in the multiple spaces; generating multiple first controls corresponding to the multiple second keywords and multiple second controls corresponding to the multiple third keywords through front-end page generation rules; determining the second position relationship between the multiple first controls and the multiple second controls based on the first position relationship between the multiple spaces and the multiple IoT devices; and generating a first control page of the terminal device based on the multiple first controls, the multiple second controls and the second position relationship.
[0044] Optionally, assuming that multiple first keywords are master bedroom, second bedroom, living room, bathroom, ceiling lamp, grille lamp, spotlight, downlight, multiple second keywords are determined to be master bedroom, second bedroom, living room, bathroom, and multiple third keywords are ceiling lamp, grille lamp, spotlight, downlight. Multiple first controls and multiple second controls corresponding to the above multiple second keywords and multiple third keywords are rendered through the front-end page generation rules. The positional relationship between the multiple first controls and the multiple second controls is determined according to the actual spatial layout. For example, if the living room has a ceiling lamp, two grille lamps, and two embedded spotlights, the first control page is generated according to the first control corresponding to the living room and the second controls corresponding to the ceiling lamp, grille lamp, and spotlight.
[0045] In an exemplary embodiment, after generating the first control page of the terminal device according to the multiple first controls, the multiple second controls and the second position relationship, the method further includes: receiving an adjustment instruction sent by the target object, and obtaining a fourth keyword corresponding to the adjustment instruction through the Internet of Things interface, wherein the adjustment instruction is used to add or delete the fourth keyword in the family metadata; according to the adjustment instruction, adjusting the control corresponding to the fourth keyword on the first control page through the control interface of the front-end page, wherein the front-end page is used to display the first control page.
[0046] Optionally, assuming that the user says to the terminal device: Use a crystal chandelier to replace the ceiling lamp in the living room, the large language model parses the above adjustment instructions, determines that the fourth keyword is a crystal chandelier and a ceiling lamp, and determines that the space that needs to be adjusted is the living room. Assuming that the unadjusted family metadata is: Living room: Ceiling lamp {quantity: 1}, Grille lamp {quantity: 2}, Spotlight {quantity: 2}, when the keyword chandelier exists in the database, the adjusted family metadata is: Living room: Chandelier {quantity: 1}, Grille lamp {quantity: 2}, Spotlight {quantity: 2}, wherein the adjustment process of the family metadata will generate corresponding operation records. According to the operation record, the control corresponding to the ceiling lamp is deleted, and the control corresponding to the chandelier is added, thereby updating the first control page corresponding to the living room.
[0047] In an exemplary embodiment, before parsing the first text or the second text through a large language model and obtaining multiple first keywords in the first text or the second text, the method also includes: changing the format of the third text to obtain a third text in a standard format, wherein the third text is the first text or the second text; performing noise filtering on the third text in the standard format to obtain a clarified third text; and extracting multiple fifth keywords from the clarified third text, wherein the multiple fifth keywords include at least one of the following: the multiple first keywords, a sixth keyword, and the sixth keyword is used to indicate the time or action corresponding to the third text.
[0048] Suppose a user enters the following text description into a terminal device: "At nine o'clock in the evening, I need to turn off the lights in the living room." The terminal device first receives the original text entered by the user, i.e., the third text. However, the actual user input may contain non-standard formats, such as multiple ways of writing dates and times ("nine o'clock in the evening," "9pm," "21:00," etc.), or the use of non-technical terms (such as "I need"). Therefore, the terminal device needs to change the format of the third text and convert all non-standard expressions into a standard format. For example, "nine o'clock in the evening" can be uniformly converted to "21:00" to facilitate subsequent parsing and the execution of time-triggered operations.
[0049] Furthermore, user input may contain noise irrelevant to control operations, such as emotional expressions, interjections, and redundant descriptions. Noise filtering can remove or ignore these irrelevant words, making the text clearer and reducing interference during parsing. For example, expressions like "I need" can be filtered out, retaining only the directly relevant description "turn off the living room lights."
[0050] After formatting and filtering the text, the terminal device extracts multiple fifth keywords from the clarified third text. The fifth keywords include the first keyword (i.e., information related to a specific device or space, such as "living room" or "lamp") and the sixth keyword (a word indicating time or action, such as "9 o'clock in the evening" or "close").
[0051] In an exemplary embodiment, after constructing family metadata based on the multiple first keywords, the method further includes: obtaining a scene control instruction issued by the target object from the first text or the second text, wherein the scene control instruction is used to control the target Internet of Things device to perform a target action; determining a target space from the multiple spaces based on the scene control instruction, and determining a second control page corresponding to the target space; and displaying the control instruction on the second control page.
[0052] Alternatively, in one embodiment, suppose a user issues a scene control instruction to a terminal device: turn on the ceiling light in the living room at 9:00 AM every day. The terminal device determines from the scene control instruction that the target space is the living room, determines the second control page corresponding to the living room, and displays the user's control instruction on the second control page. The user can modify the control instruction on the second control page.
[0053] In another embodiment, suppose the user says to the terminal device: Set goodnight mode and turn off the bedroom lights. The large language model parses the control instruction to determine the scenario mode (goodnight mode), the target space (bedroom), and the target IoT device (all lights) in the target space. After the analysis is completed, the terminal device will determine the second control page corresponding to the bedroom and display the scenario mode on the second control page. In subsequent use, the user only needs to send a voice instruction corresponding to the scenario mode to the terminal device to implement the control instruction corresponding to the scenario mode. For example, when there is a goodnight mode in the second control page, the user only needs to say the keyword with the goodnight mode to turn off the bedroom lights.
[0054] In an exemplary embodiment, after constructing family metadata based on the multiple first keywords, the method also includes: obtaining the status of multiple IoT devices in the multiple spaces through the IoT interface; when there is a first IoT device among the multiple IoT devices and the status of the first IoT device is stopped, deleting the third keyword corresponding to the first IoT device from the family metadata; when a second IoT device is added to the multiple spaces, updating the family metadata according to the seventh keyword corresponding to the second IoT device.
[0055] The terminal device continuously obtains the operating status of these IoT devices through the IoT devices deployed in various spaces of the home (such as smart light bulbs, air conditioners, security cameras, etc.) and their corresponding IoT interfaces. The operating status may include whether the device is online, whether it is in working state, power level, fault status, etc. When the IoT interface detects that a device in a certain space, such as a smart air conditioner in the living room (i.e., the first IoT device), its status changes to "out of use" (possibly because the device is physically removed, has not been connected for a long time, or the user manually sets it to offline), the terminal device will automatically determine the device corresponding to the living room from the home metadata and delete the third keyword related to the air conditioner. When the user adds a new device to the home, such as adding a smart security camera (the second IoT device) to the master bedroom, the terminal device matches the camera with multiple keywords in the database. If the match is successful, the terminal device will update the home metadata based on the seventh keyword of the second IoT device.
[0056] In an exemplary embodiment, after parsing the first text or the second text through a large language model and obtaining multiple first keywords in the first text or the second text, the method further includes: when there is an eighth keyword among the multiple first keywords that fails to match multiple keywords in the database of the cloud server, receiving an error message sent by the cloud server, wherein the error message includes the reason why the eighth keyword fails to match the multiple keywords.
[0057] Suppose a user says to a terminal device, "Please turn on the smart fan on the balcony." Here, "balcony" and "smart fan" are the first keywords extracted from the user's command. The large language model compares these keywords with the cloud server's database to verify whether they correspond to known spaces or devices. The cloud server's database contains all registered device information and space names, ensuring that the user's command is correctly executed. If no matching keyword for "balcony" or "smart fan" (i.e., the eighth keyword) is found in the database, this means that the device or space information has not yet been registered. In this case, the terminal device cannot execute the control operation specified in the user's command. To help the user understand the problem, the cloud server generates and sends an error message to the user's terminal device. The error message includes the specific reason for the matching failure, such as "No matching space or device information found for balcony or smart fan." Please check whether the device is registered and named correctly, or confirm whether the space name is correct. After receiving the error message, the user can check whether "balcony" is defined on the terminal device and whether "smart fan" has been added to the device list in the home metadata, and ensure that the device name matches the naming rules in the database.
[0058] In order to better understand the process of the above-mentioned family metadata construction method, the above-mentioned family metadata construction method is described below in combination with an optional embodiment, but is not used to limit the technical solution of the embodiment of this application.
[0059] Figure 3 This is a structural block diagram of a system for generating a control page according to an embodiment of the present application. Figure 3 As shown, the specific steps include:
[0060] Step S302: obtaining a voice instruction of the target object or a text (ie, a second text) directly input by the target object through a user voice / text input module.
[0061] Step S304: The speech recognition / text preprocessing module converts the speech instruction acquired in step S302 into a first text and performs text preprocessing on the first or second text. The text preprocessing includes, but is not limited to, formatting changes and noise filtering. Finally, keywords are extracted from the preprocessed first or second text.
[0062] Step S306: The keywords extracted in step S304 are parsed by the large language model parsing module to determine a plurality of first keywords corresponding one-to-one to a plurality of spaces in the target area where the target object is located or a plurality of IoT devices in the plurality of spaces.
[0063] Step S308: The decision and matching module matches the multiple first keywords in step S306 with multiple keywords in the database of the cloud server. If the target keyword that matches the first keyword exists in the database, the process proceeds to step 310. If some of the multiple first keywords fail to match multiple keywords in the database, this means that these keywords are not registered in the database. In this case, the cloud server generates and sends an error message to the user's terminal device. The error message will include the specific reason for the match failure, such as an unrecognized control name or a product match error.
[0064] Step S310: Constructing family metadata based on multiple first keywords through the metadata generation module. Optionally, assume that the user says to the terminal device: My home has two bedrooms and one living room, a master bedroom, a second bedroom, a living room, and a bathroom. The living room has a ceiling lamp, two grille lamps, and two embedded spotlights. The second bedroom has four downlights. The master bedroom has a ceiling lamp and four spotlights. The bathroom has a ceiling lamp. The large language model parsing module of step S306 can determine that the multiple first keywords include master bedroom, second bedroom, living room, bathroom, ceiling lamp, grille lamp, spotlight, and downlight. At the same time, the large language model will also understand the relationship between the above multiple first keywords during the parsing process, for example, which devices belong to which room and how many devices are in each room, and extract this information through a dedicated prompt or instruction set. The final household metadata obtained is: Living room: Ceiling light {quantity: 1}, Grille light {quantity: 2}, Spotlight {quantity: 2}; Second bedroom: Downlight {quantity: 4}; Master bedroom: Ceiling light {quantity: 1}, Spotlight {quantity: 4}; Bathroom: Ceiling light {quantity: 1}.
[0065] At the same time, the terminal device continuously obtains the operating status of these IoT devices through the IoT devices deployed in various spaces in the home and their corresponding IoT interfaces. When the IoT interface detects that a device in a certain space, such as a smart air conditioner in the living room (i.e., the first IoT device), changes its status to stopped, the terminal device will automatically determine the device corresponding to the living room from the home metadata and delete the keywords related to the air conditioner. When the user adds a new device to the home, such as adding a smart security camera (the second IoT device) to the master bedroom, the terminal device matches the camera with multiple keywords in the database. If the match is successful, the terminal device will update the home metadata based on the keywords of the second IoT device.
[0066] Step S312: Determine, through the control page generation module, multiple second keywords corresponding to the multiple spaces contained in the multiple first keywords, and multiple third keywords corresponding to the multiple IoT devices in the multiple spaces. Generate multiple first controls corresponding to the multiple second keywords and multiple second controls corresponding to the multiple third keywords using the front-end page generation rules. Determine the second positional relationship between the multiple first controls and the multiple second controls based on the first positional relationship between the multiple spaces and the multiple IoT devices. Generate a control page for the terminal device based on the multiple first controls, the multiple second controls, and the second positional relationship.
[0067] If a user needs to adjust a control page, they can send an adjustment command to the terminal device. For example, to replace the living room ceiling light with a chandelier, the large language model will determine the control page for the living room based on the adjustment command and retrieve the operation record for the household metadata. This operation record records the changes to the living room devices in the household metadata. For example, the unadjusted household metadata is: Living Room: Ceiling Light {Quantity: 1}, Grille Light {Quantity: 2}, Spotlight {Quantity: 2}. If the keyword "chandelier" exists in the database, the adjusted household metadata will be: Living Room: Chandelier {Quantity: 1}, Grille Light {Quantity: 2}, Spotlight {Quantity: 2}. Based on the operation record, the terminal device will delete the control corresponding to the ceiling light and add the control corresponding to the chandelier, thereby updating the control page for the living room.
[0068] When a user issues a scene control command, the target space corresponding to the scene control command is determined, the control page corresponding to the target space is determined, and the scene control command is displayed on the control page corresponding to the target space. Alternatively, in one embodiment, suppose a user issues a scene control command to a terminal device: "Turn on the living room ceiling light every morning at 9:00 AM." The terminal device determines from the scene control command that the target space is the living room, determines the control page corresponding to the living room, and displays the user's control command on the living room control page. In another embodiment, suppose a user says to the terminal device: "Set goodnight mode to turn off the bedroom lights." The large language model parses the control command to determine the scene mode (goodnight mode), the target space (bedroom), and the target IoT devices within the target space (all lights). After parsing, the terminal device determines the control page corresponding to the bedroom and displays the scene mode on the bedroom control page. In subsequent use, the user only needs to issue a voice command corresponding to the scene mode to the terminal device to implement the control command corresponding to the scene mode. For example, if the bedroom control page contains a goodnight mode, the user can simply say the keyword containing the goodnight mode to turn off the bedroom lights.
[0069] Figure 4 This is a first schematic diagram of a method for constructing family metadata according to an embodiment of the present application. Figure 4 As shown, specifically including the following:
[0070] Alternatively, assume that a user issues a voice command to a terminal device: "My home has two bedrooms and a living room: a master bedroom, a second bedroom, a living room, and a bathroom. The living room has one ceiling light, two grille lights, and two recessed spotlights. The second bedroom has four downlights. The master bedroom has one ceiling light and four spotlights. The bathroom has one ceiling light." The user's voice command is converted into a first text using speech recognition technology, or the user directly enters a second text in the input box. After entering the first or second text, the user clicks "Generate Design Plan." The large language model then begins parsing the first or second text to identify multiple first keywords within the text.
[0071] Figure 5 This is a second schematic diagram of a method for constructing family metadata according to an embodiment of the present application. Figure 5 As shown, specifically including the following:
[0072] Optionally, assuming that the multiple first keywords are master bedroom, second bedroom, living room, bathroom, ceiling light, grille light, spotlight, downlight, and the multiple second keywords are determined to be master bedroom, second bedroom, living room, bathroom, and the multiple third keywords are ceiling light, grille light, spotlight, downlight. The multiple first controls and multiple second controls corresponding to the multiple second keywords and multiple third keywords are rendered through the front-end page generation rules. The positional relationship between the multiple first controls and the multiple second controls is determined according to the actual spatial layout. For example, assuming that the master bedroom has four downlights, the first controls corresponding to the second keywords such as master bedroom, second bedroom, living room, bathroom, etc. are displayed on the control page of the master bedroom, and according to the corresponding relationship between the master bedroom and the downlights, the second controls of the downlights (i.e. downlight 1, downlight 2, downlight 3 and downlight 4 in the figure) are obtained from the database of the cloud server, and the second controls of the four downlights are displayed on the control page of the master bedroom.
[0073] Furthermore, the control page can also display user-defined scenarios and automation commands. In one optional embodiment, suppose a user issues a voice command to a terminal device: "Set Home Mode to turn on the living room lights," or "Set Away Mode to turn off all lights in all rooms." The large language model parses the control command, determining the scenario mode, the target space, and the target IoT device within the target space. After parsing, the terminal device determines the control page corresponding to the target space and displays the scenario mode on the control page. In subsequent use, the user simply issues a voice command corresponding to the scenario mode to the terminal device to execute the control command corresponding to that scenario mode. For example, if the control page displays Home Mode, the user simply speaks a keyword containing Home Mode to turn on the living room lights. In another optional embodiment, suppose a user issues a scenario control command to the terminal device: "Turn on the master bedroom lights every morning at 9:00 AM." The terminal device determines from the scenario control command that the target space is the master bedroom, determines the control page corresponding to the master bedroom, and displays the user's control command on the control page.
[0074] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0075] Figure 6 is a structural block diagram of a device for constructing family metadata according to an embodiment of the present application; Figure 6 Shown, including:
[0076] An acquisition module 62 is configured to acquire a voice command sent by a target object and convert the voice command into a first text using a voice recognition technology, or to acquire a second text sent by the target object;
[0077] a parsing module 64, configured to parse the first text or the second text using a large language model to obtain a plurality of first keywords in the first text or the second text, wherein the plurality of first keywords correspond one-to-one to a plurality of spaces in a target area where the target object is located or a plurality of IoT devices in the plurality of spaces;
[0078] Construction module 66 is used to construct family metadata based on the multiple first keywords when the multiple first keywords successfully match the target keywords in the database of the cloud server, wherein the family metadata is at least used to characterize the positional relationship between the multiple IoT devices and the multiple spaces.
[0079] The above-mentioned device obtains a voice command issued by a target object and converts the voice command into a first text using speech recognition technology, or obtains a second text sent by the target object. The first text or the second text is parsed using a large language model to obtain multiple first keywords in the first text or the second text, wherein the multiple first keywords correspond one-to-one with multiple spaces in the target area where the target object is located or multiple IoT devices in the multiple spaces. When the multiple first keywords successfully match the target keywords in the database of the cloud server, family metadata is constructed based on the multiple first keywords, wherein the family metadata is used to at least characterize the positional relationship between the multiple IoT devices and the multiple spaces. The above-mentioned method solves the problem in the related art of manually inputting family metadata, which is not only time-consuming and labor-intensive, but also prone to errors in the family metadata.
[0080] In an exemplary embodiment, construction module 66 is also used to determine multiple second keywords included in the multiple first keywords corresponding to the multiple spaces, and multiple third keywords corresponding to the multiple IoT devices in the multiple spaces; generate multiple first controls corresponding to the multiple second keywords and multiple second controls corresponding to the multiple third keywords through front-end page generation rules; determine the second position relationship between the multiple first controls and the multiple second controls based on the first position relationship between the multiple spaces and the multiple IoT devices; and generate a first control page for the terminal device based on the multiple first controls, the multiple second controls and the second position relationship.
[0081] In an exemplary embodiment, the construction module 66 is also used to receive an adjustment instruction sent by the target object, and obtain a fourth keyword corresponding to the adjustment instruction through the Internet of Things interface, wherein the adjustment instruction is used to add or delete the fourth keyword in the family metadata; according to the adjustment instruction, the control corresponding to the fourth keyword is adjusted on the first control page through the control interface of the front-end page, wherein the front-end page is used to display the first control page.
[0082] In an exemplary embodiment, the parsing module 64 is further used to change the format of the third text to obtain a third text in a standard format, wherein the third text is the first text or the second text; perform noise filtering on the third text in the standard format to obtain a clarified third text; and extract multiple fifth keywords from the clarified third text, wherein the multiple fifth keywords include at least one of the following: the multiple first keywords, a sixth keyword, and the sixth keyword is used to indicate the time or action corresponding to the third text.
[0083] In an exemplary embodiment, construction module 66 is further used to obtain a scene control instruction issued by the target object from the first text or the second text, wherein the scene control instruction is used to control the target IoT device to perform a target action; determine a target space from the multiple spaces according to the scene control instruction, and determine a second control page corresponding to the target space; and display the control instruction on the second control page.
[0084] In an exemplary embodiment, building module 66 is also used to obtain the status of multiple IoT devices in the multiple spaces through the IoT interface; when there is a first IoT device among the multiple IoT devices and the status is stopped, the third keyword corresponding to the first IoT device is deleted from the family metadata; when a second IoT device is added to the multiple spaces, the family metadata is updated according to the seventh keyword corresponding to the second IoT device.
[0085] In an exemplary embodiment, the parsing module 64 is also used to receive an error message sent by the cloud server when there is an eighth keyword among the multiple first keywords that fails to match multiple keywords in the database of the cloud server, wherein the error message includes the reason why the eighth keyword fails to match the multiple keywords.
[0086] In this embodiment, a method for generating a control page is also provided, which is applied to a terminal device. Figure 7 1 is a flow chart of a method for generating a control page according to an embodiment of the present application, the flow comprising the following steps:
[0087] Step S702: Acquire a voice command sent by a target object, and convert the voice command into a first text using a voice recognition technology, or acquire a second text sent by the target object;
[0088] Step S704: Parsing the first text or the second text using a large language model to obtain multiple first keywords in the first text or the second text, wherein the multiple first keywords correspond one-to-one to multiple spaces in the target area where the target object is located or multiple IoT devices in the multiple spaces;
[0089] Step S706: If the plurality of first keywords successfully match the target keywords in the database of the cloud server, constructing family metadata based on the plurality of first keywords, wherein the family metadata is used to at least characterize the positional relationship between the plurality of IoT devices and the plurality of spaces;
[0090] Step S708: Generate a target control corresponding to the family metadata through the front-end page generation rules, and generate a first control page based on the target control.
[0091] Through the above steps, first, a voice command issued by the target object is obtained, and the voice command is converted into a first text through voice recognition technology, or a second text sent by the target object is obtained; secondly, the first text or the second text is parsed through a large language model to obtain multiple first keywords in the first text or the second text, wherein the multiple first keywords correspond one-to-one to multiple spaces in the target area where the target object is located or multiple IoT devices in the multiple spaces; when the multiple first keywords successfully match the target keywords in the database of the cloud server, family metadata is constructed based on the multiple first keywords, wherein the family metadata is at least used to characterize the positional relationship between the multiple IoT devices and the multiple spaces; finally, a target control corresponding to the family metadata is generated through the front-end page generation rules, and a first control page is generated based on the target control. The above method solves the problem in the related art that the target object manually generates a control page, which increases the complexity of manual operation and has low control page generation efficiency.
[0092] Figure 8 This is a structural block diagram of a device for generating a control page according to an embodiment of the present application. Figure 8 Shown, including:
[0093] An acquisition module 82 is configured to acquire a voice command sent by a target object and convert the voice command into a first text using a voice recognition technology, or to acquire a second text sent by the target object;
[0094] a parsing module 84, configured to parse the first text or the second text using a large language model to obtain a plurality of first keywords in the first text or the second text, wherein the plurality of first keywords correspond one-to-one to a plurality of spaces in a target area where the target object is located or a plurality of IoT devices in the plurality of spaces;
[0095] a construction module 86 configured to construct family metadata based on the plurality of first keywords if the plurality of first keywords successfully match target keywords in a database of the cloud server, wherein the family metadata is used to at least characterize a positional relationship between the plurality of IoT devices and the plurality of spaces;
[0096] The generation module 88 is used to generate a target control corresponding to the family metadata through a front-end page generation rule, and generate a first control page according to the target control.
[0097] The above-mentioned device first obtains a voice command issued by a target object and converts the voice command into a first text through voice recognition technology, or obtains a second text sent by the target object; secondly, the first text or the second text is parsed through a large language model to obtain multiple first keywords in the first text or the second text, wherein the multiple first keywords correspond one-to-one to multiple spaces in the target area where the target object is located or multiple IoT devices in the multiple spaces; if the multiple first keywords successfully match the target keywords in the database of the cloud server, family metadata is constructed based on the multiple first keywords, wherein the family metadata is used to at least characterize the positional relationship between the multiple IoT devices and the multiple spaces; finally, a target control corresponding to the family metadata is generated through a front-end page generation rule, and a first control page is generated based on the target control. The above-mentioned method solves the problem in the related art that the target object manually generates a control page, which increases the complexity of manual operation and has low control page generation efficiency.
[0098] An embodiment of the present application further provides a storage medium, which includes a stored program, wherein the program executes any of the above methods when it is run.
[0099] Optionally, in this embodiment, the storage medium may be configured to store program codes for executing the following steps:
[0100] S11, obtaining a voice command issued by a target object, and converting the voice command into a first text using a voice recognition technology, or obtaining a second text sent by the target object;
[0101] S12, parsing the first text or the second text using a large language model to obtain multiple first keywords in the first text or the second text, wherein the multiple first keywords correspond one-to-one to multiple spaces in the target area where the target object is located or multiple IoT devices in the multiple spaces;
[0102] S13, when the multiple first keywords are successfully matched with the target keywords in the database of the cloud server, constructing family metadata based on the multiple first keywords, wherein the family metadata is at least used to characterize the positional relationship between the multiple IoT devices and the multiple spaces.
[0103] Optionally, in this embodiment, the storage medium may also be configured to store program codes for executing the following steps:
[0104] S21, obtaining a voice command issued by a target object, and converting the voice command into a first text using a voice recognition technology, or obtaining a second text sent by the target object;
[0105] S22, parsing the first text or the second text using a large language model to obtain a plurality of first keywords in the first text or the second text, wherein the plurality of first keywords correspond one-to-one to a plurality of spaces in a target area where the target object is located or a plurality of IoT devices in the plurality of spaces;
[0106] S23, if the multiple first keywords successfully match the target keywords in the database of the cloud server, constructing family metadata based on the multiple first keywords, wherein the family metadata is used to at least characterize the positional relationship between the multiple IoT devices and the multiple spaces;
[0107] S24, generating a target control corresponding to the family metadata through a front-end page generation rule, and generating a first control page according to the target control.
[0108] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0109] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0110] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0111] S11, obtaining a voice command issued by a target object, and converting the voice command into a first text using a voice recognition technology, or obtaining a second text sent by the target object;
[0112] S12, parsing the first text or the second text using a large language model to obtain multiple first keywords in the first text or the second text, wherein the multiple first keywords correspond one-to-one to multiple spaces in the target area where the target object is located or multiple IoT devices in the multiple spaces;
[0113] S13, when the multiple first keywords are successfully matched with the target keywords in the database of the cloud server, constructing family metadata based on the multiple first keywords, wherein the family metadata is at least used to characterize the positional relationship between the multiple IoT devices and the multiple spaces.
[0114] Optionally, in this embodiment, the processor may be further configured to execute the following steps through a computer program:
[0115] S21, obtaining a voice command issued by a target object, and converting the voice command into a first text using a voice recognition technology, or obtaining a second text sent by the target object;
[0116] S22, parsing the first text or the second text using a large language model to obtain a plurality of first keywords in the first text or the second text, wherein the plurality of first keywords correspond one-to-one to a plurality of spaces in a target area where the target object is located or a plurality of IoT devices in the plurality of spaces;
[0117] S23, if the multiple first keywords successfully match the target keywords in the database of the cloud server, constructing family metadata based on the multiple first keywords, wherein the family metadata is used to at least characterize the positional relationship between the multiple IoT devices and the multiple spaces;
[0118] S24, generating a target control corresponding to the family metadata through a front-end page generation rule, and generating a first control page according to the target control.
[0119] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store program codes.
[0120] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0121] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices. Alternatively, they can be implemented using program code executable by the computing device, so that they can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than herein, or they can be made into separate integrated circuit modules, or multiple modules or steps can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.
[0122] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A method for constructing family metadata, characterized in that: Applied to terminal equipment, including: Acquire a voice command issued by a target object, and convert the voice command into a first text using a voice recognition technology, or acquire a second text sent by the target object; Parsing the first text or the second text using a large language model to obtain a plurality of first keywords in the first text or the second text, wherein the plurality of first keywords correspond one-to-one to a plurality of spaces in a target area where the target object is located or a plurality of IoT devices in the plurality of spaces; When the multiple first keywords successfully match the target keywords in the database of the cloud server, family metadata is constructed based on the multiple first keywords, wherein the family metadata is at least used to characterize the positional relationship between the multiple IoT devices and the multiple spaces.
2. The method for constructing family metadata according to claim 1, characterized in that: After constructing the family metadata according to the plurality of first keywords, the method further includes: Determining a plurality of second keywords included in the plurality of first keywords and corresponding to the plurality of spaces, and a plurality of third keywords corresponding to the plurality of IoT devices in the plurality of spaces; Generate a plurality of first controls corresponding to the plurality of second keywords and a plurality of second controls corresponding to the plurality of third keywords through a front-end page generation rule; Determine a second positional relationship between the plurality of first controls and the plurality of second controls according to the first positional relationship between the plurality of spaces and the plurality of IoT devices; A first control page of the terminal device is generated according to the multiple first controls, the multiple second controls, and the second position relationship.
3. The method for constructing family metadata according to claim 2, characterized in that: After generating the first control page of the terminal device according to the plurality of first controls, the plurality of second controls, and the second position relationship, the method further includes: receiving an adjustment instruction sent by the target object, and obtaining a fourth keyword corresponding to the adjustment instruction through an Internet of Things interface, wherein the adjustment instruction is used to add or delete the fourth keyword in the family metadata; According to the adjustment instruction, the control corresponding to the fourth keyword is adjusted on the first control page through the control interface of the front-end page, wherein the front-end page is used to display the first control page.
4. The method for constructing family metadata according to claim 1, characterized in that: Before parsing the first text or the second text using a large language model to obtain a plurality of first keywords in the first text or the second text, the method further includes: performing formatting changes on a third text to obtain a third text in a standard format, wherein the third text is the first text or the second text; performing noise filtering on the third text in the standard format to obtain a clarified third text; A plurality of fifth keywords are extracted from the clarified third text, wherein the plurality of fifth keywords include at least one of the following: the plurality of first keywords and a sixth keyword, and the sixth keyword is used to indicate a time or action corresponding to the third text.
5. The method for constructing family metadata according to claim 1, characterized in that: After constructing the family metadata according to the plurality of first keywords, the method further includes: Obtaining a scene control instruction issued by the target object from the first text or the second text, wherein the scene control instruction is used to control a target IoT device to perform a target action; determining a target space from the multiple spaces according to the scene control instruction, and determining a second control page corresponding to the target space; The control instruction is displayed on the second control page.
6. The method for constructing family metadata according to claim 2, characterized in that: After constructing the family metadata according to the plurality of first keywords, the method further includes: Obtaining, through an IoT interface, the status of multiple IoT devices in the multiple spaces; When a first IoT device among the plurality of IoT devices is in a stopped-use state, deleting a third keyword corresponding to the first IoT device from the family metadata; When a second Internet of Things device is newly added to the multiple spaces, the family metadata is updated according to the seventh keyword corresponding to the second Internet of Things device.
7. The method for constructing family metadata according to claim 1, characterized in that: After parsing the first text or the second text using a large language model to obtain a plurality of first keywords in the first text or the second text, the method further includes: In the event that an eighth keyword among the multiple first keywords fails to match multiple keywords in the database of the cloud server, an error message sent by the cloud server is received, wherein the error message includes a reason why the eighth keyword fails to match the multiple keywords.
8. A device for constructing family metadata, characterized in that: include: An acquisition module, configured to acquire a voice instruction sent by a target object and convert the voice instruction into a first text using a voice recognition technology, or to acquire a second text sent by the target object; a parsing module, configured to parse the first text or the second text using a large language model to obtain a plurality of first keywords in the first text or the second text, wherein the plurality of first keywords correspond one-to-one to a plurality of spaces in a target area where the target object is located or a plurality of IoT devices in the plurality of spaces; A construction module is used to construct family metadata based on the multiple first keywords when the multiple first keywords are successfully matched with the target keywords in the database of the cloud server, wherein the family metadata is at least used to characterize the positional relationship between the multiple Internet of Things devices and the multiple spaces.
9. A method for generating a control page, characterized in that: Applied to terminal equipment, including: Acquire a voice command issued by a target object, and convert the voice command into a first text using a voice recognition technology, or acquire a second text sent by the target object; Parsing the first text or the second text using a large language model to obtain a plurality of first keywords in the first text or the second text, wherein the plurality of first keywords correspond one-to-one to a plurality of spaces in a target area where the target object is located or a plurality of IoT devices in the plurality of spaces; If the plurality of first keywords successfully match the target keywords in the database of the cloud server, constructing family metadata based on the plurality of first keywords, wherein the family metadata is at least used to characterize the positional relationship between the plurality of IoT devices and the plurality of spaces; A target control corresponding to the family metadata is generated through the front-end page generation rules, and a first control page is generated according to the target control.
10. A device for generating a control page, characterized in that: include: An acquisition module, configured to acquire a voice instruction sent by a target object and convert the voice instruction into a first text using a voice recognition technology, or to acquire a second text sent by the target object; a parsing module, configured to parse the first text or the second text using a large language model to obtain a plurality of first keywords in the first text or the second text, wherein the plurality of first keywords correspond one-to-one to a plurality of spaces in a target area where the target object is located or a plurality of IoT devices in the plurality of spaces; a construction module, configured to construct family metadata based on the plurality of first keywords when the plurality of first keywords successfully match target keywords in a database of the cloud server, wherein the family metadata is used to at least characterize a positional relationship between the plurality of IoT devices and the plurality of spaces; A generation module is used to generate a target control corresponding to the family metadata through a front-end page generation rule, and generate a first control page according to the target control.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 7 or 9 when executed.
12. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 or 9 through the computer program.