Control method and device of spherical camera and electronic equipment

By receiving user input data and generating control instructions using the large language model in the intelligent agent model, the problem of manual adjustment of spherical cameras is solved, and efficient and intelligent monitoring task execution and autonomous control are achieved.

CN120751237APending Publication Date: 2025-10-03CHINA TOWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510896181.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The existing spherical camera control method requires manual adjustment according to the operating manual and camera model, resulting in low work efficiency.

Method used

By receiving text or voice data input by the target user, the large language model in the intelligent agent model is used to parse the task to be processed, generate and execute control instructions, and generate output results including text, images or video information by updating the control instructions multiple times until the task is completed.

Benefits of technology

It achieves efficient identification of user needs, automatically generates precise control instructions, completes complex tasks such as target detection and tracking, improves monitoring efficiency and accuracy, reduces manual intervention, and adapts to the ever-changing technological environment and user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751237A_ABST
    Figure CN120751237A_ABST
Patent Text Reader

Abstract

The invention discloses a spherical camera control method and device and electronic equipment, the method is applied to the field of artificial intelligence, and the method comprises the following steps: receiving target data input by a target user; a to-be-processed task in the target data is analyzed through a large language model in the agent model, a control instruction is generated and executed, an execution result is obtained, the control instruction is an instruction for controlling the multiple dome cameras in the camera group according to codes in the code library, and the execution result is obtained. The code library comprises codes for calling a plurality of deep learning models in the agent model; the execution result is input into the agent model, the control instruction is updated multiple times through the agent model, the control instruction updated each time is executed, the execution result is updated till the task to be processed is completed, and an output result is generated. Through application of the method and the device, the problem that the working efficiency of the camera is relatively low due to the fact that repeated adjustment needs to be manually performed according to an operation manual and a camera model when the spherical camera is controlled to execute the shooting task in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and more specifically, to a control method, device, and electronic device for a spherical camera. Background Art

[0002] With the development of artificial intelligence (AI), network cameras have evolved from traditional infrared / black and white / color pixel imaging to today's 4K / 8K / 1080P high-definition image quality. Their use cases have expanded from static images to recording, streaming, and live broadcasting. They have evolved from single-use infrared / color / black and white capture to full-time, full-color, and full-spectrum capture. They have evolved from standard image quality formats to custom formats, from hardware upgrades to algorithm optimization, from cloud platform connectivity to end-device access, and from simple data collection and transmission to intelligent analysis applications. Network cameras have also evolved from their early applications of simple monitoring and recording to today's diverse applications, including real-time image transmission, intelligent security, and edge computing.

[0003] Current intelligent control of network cameras primarily relies on standalone control of traditional network cameras, enabling autonomous control through pre-set hardware device control commands. However, with the increasing complexity of smart security environments, network camera control requirements have become more diverse and personalized. Existing camera control logic relies on a human operator observing video content and then issuing corresponding commands to the camera. Users must follow the operating manual for each camera model, and the user interface is inconsistent, resulting in low camera efficiency.

[0004] Regarding the problem in related technologies where controlling a spherical camera to perform shooting tasks requires repeated manual adjustments based on the operating manual and camera model, resulting in low camera operating efficiency, no effective solution has yet been proposed. Summary of the Invention

[0005] The main purpose of this application is to provide a control method, device and electronic equipment for a spherical camera to solve the problem in the related art that when controlling a spherical camera to perform a shooting task, manual adjustments need to be made repeatedly according to the operating manual and camera model, resulting in low camera working efficiency.

[0006] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a control method for a spherical camera is provided, which includes: receiving target data input by a target user, wherein the target data includes at least one of the following: text data, voice data; parsing the pending tasks in the target data through a large language model in an intelligent agent model, generating and executing control instructions, and obtaining execution results, wherein the control instructions are instructions for controlling multiple spherical cameras in a camera group based on codes in a code library, and the code library contains codes for calling multiple deep learning models in the intelligent agent model; inputting the execution results into the intelligent agent model, updating the control instructions multiple times through the intelligent agent model, executing the updated control instructions each time, and updating the execution results until the pending tasks are completed, and generating output results, wherein the output results include at least one of the following: text information, image information, and video information.

[0007] Furthermore, the execution result is input into the intelligent agent model, the control instruction is updated multiple times through the intelligent agent model, the updated control instruction is executed each time, and the execution result is updated until the pending task is completed to generate an output result, including: inputting the execution result into the large language model for the Nth time, analyzing and making decisions on the target data through the large language model, and obtaining the Nth updated control instruction; calling the code in the code library according to the Nth updated control instruction, updating the execution result, and obtaining the Nth execution result corresponding to the Nth updated control instruction; judging whether to continue to update the control instruction through the large language model according to the pending task and the Nth execution result; in the case of updating the control instruction, storing the Nth execution result, and inputting the Nth execution result into the large language model for the N+1th time; in the case of not updating the control instruction, generating the output result according to the execution result corresponding to each execution of the control instruction.

[0008] Furthermore, the large language model in the intelligent agent model is used to parse the tasks to be processed in the target data, generate and execute control instructions, and obtain execution results, including: thinking about the tasks to be processed through the large language model to obtain the output results of the large language model; converting the output results into the control instructions according to a preset output format, wherein the control instructions include control operations, and the control instructions include at least the following information: the tasks to be processed, the execution operations, the input parameter information of the execution operations, and the expected execution results of the execution operations; executing the control instructions to determine the execution results.

[0009] Furthermore, the large language model in the intelligent agent model is used to parse the pending tasks in the target data, generate and execute control instructions, and obtain execution results, including: when the large language model and the multiple deep learning models are deployed on a local computing device, the target data is sent to the local computing device, and the pending tasks in the target data are parsed by the large language model deployed in the local computing device, and the control instructions are generated and executed to obtain the execution results, wherein the computing power of the local computing device is greater than the preset computing power, and the control instructions are used to call the multiple deep learning models in the local computing device to perform tasks; or, when the target model is deployed on a cloud computing device, the pending tasks in the target data are parsed by the large language model, and the control instructions are generated and executed, and the control instructions are sent to the cloud computing device, wherein the target model is a model in which the number of model parameters of the multiple deep learning models is greater than the preset number, and the control instructions are used to call the target model in the cloud computing device to perform tasks; and the execution results sent by the cloud computing device are received.

[0010] Furthermore, before parsing the pending tasks in the target data through the large language model in the intelligent agent model, the method also includes: processing the target data and the domain knowledge base according to a preset format to obtain prompt words, wherein the prompt words include at least one user intent, so that the large language model parses each user intent in the at least one user intent to obtain the pending tasks; the domain knowledge base at least includes code information corresponding to the control instruction; determining the preset identity information of the intelligent agent model based on the target data, updating the prompt words based on the preset identity information, and inputting the updated prompt words into the large language model.

[0011] Furthermore, after receiving the target data input by the target user, the method also includes: determining the time tag of the target data; creating dialogue information based on the target data and the time tag, and storing the dialogue information in a first storage space; after generating and executing a control instruction to obtain an execution result, the method also includes: determining the output time of the execution result; storing the execution result and the output time in the dialogue information of the first storage space; filtering the output information containing the detection result in the first storage space every second preset time period, and storing the output information of the detection result in the second storage space; and performing data cleaning in the second storage space according to the output time of the execution result every first preset time period.

[0012] Furthermore, the code library contains code information of multiple codes, and the multiple codes include at least one of the following: a first code for controlling the relative movement or absolute movement of the spherical camera; a second code for maintaining the preset point position of the spherical camera; a third code for collecting the status information of the spherical camera, querying the information of the spherical camera and obtaining the shooting results of the spherical camera; a fourth code for performing target detection tasks based on the multiple deep learning models; a fifth code for patrolling the environment in a preset direction or within a preset range; a sixth code for querying information; a seventh code for starting or terminating a task, and the code information includes at least: code function description information, code name, code input parameters, and input parameter information.

[0013] In order to achieve the above-mentioned purpose, according to another aspect of the present application, a control device for a spherical camera is provided, which includes: a receiving unit for receiving target data input by a target user, wherein the target data includes at least one of the following: text data, voice data; a first execution unit for parsing the pending tasks in the target data through a large language model in an intelligent agent model, generating and executing control instructions, and obtaining execution results, wherein the control instructions are instructions for controlling multiple spherical cameras in a camera group based on codes in a code library, and the code library contains codes for calling multiple deep learning models in the intelligent agent model; a second execution unit for inputting the execution results into the intelligent agent model, updating the control instructions multiple times through the intelligent agent model, executing the updated control instructions each time, and updating the execution results until the pending tasks are completed and an output result is generated, wherein the output results include at least one of the following: text information, image information, and video information.

[0014] Furthermore, the second execution unit includes: a first generation subunit, used to input the execution result into the large language model for the Nth time, analyze and make decisions on the target data through the large language model, and obtain the control instruction after the Nth update; an update subunit, used to call the code in the code library according to the control instruction after the Nth update, update the execution result, and obtain the Nth execution result corresponding to the control instruction after the Nth update; a judgment subunit, used to judge whether to continue to update the control instruction through the large language model based on the task to be processed and the Nth execution result; a storage subunit, used to store the Nth execution result when updating the control instruction, and input the Nth execution result into the large language model for the N+1th time; a second generation subunit, used to generate the output result based on the execution result corresponding to each execution of the control instruction without updating the control instruction.

[0015] Furthermore, the first execution unit includes: a third generation subunit, used to think about the task to be processed through the large language model to obtain the output result of the large language model; a conversion subunit, used to convert the output result into the control instruction according to a preset output format, wherein the control instruction includes a control operation, and the control instruction includes at least the following information: the task to be processed, the execution operation, the input parameter information of the execution operation, and the expected execution result of the execution operation; a determination subunit, used to execute the control instruction and determine the execution result.

[0016] Furthermore, the first execution unit includes: a first execution sub-unit, which is used to send the target data to the local computing device when the large language model and the multiple deep learning models are deployed on the local computing device, and parse the tasks to be processed in the target data through the large language model deployed in the local computing device, generate and execute the control instructions, and obtain the execution result, wherein the computing power of the local computing device is greater than the preset computing power, and the control instructions are used to call the multiple deep learning models in the local computing device to perform tasks; or, a second execution sub-unit, which is used to parse the tasks to be processed in the target data through the large language model when the target model is deployed on a cloud computing device, generate and execute the control instructions, and send the control instructions to the cloud computing device, wherein the target model is a model in which the number of model parameters of the multiple deep learning models is greater than the preset number, and the control instructions are used to call the target model in the cloud computing device to perform tasks; and receive the execution result sent by the cloud computing device.

[0017] Furthermore, the device also includes: a processing unit, used to process the target data and the domain knowledge base in a preset format to obtain a prompt word, wherein the prompt word includes at least one user intent, so that the large language model parses each user intent of the at least one user intent to obtain the task to be processed; the domain knowledge base at least includes code information corresponding to the control instruction; an input unit, used to determine the preset identity information of the intelligent agent model based on the target data, and update the prompt word based on the preset identity information, and input the updated prompt word into the large language model.

[0018] Furthermore, the device also includes: a first determination unit, for determining the time tag of the target data after receiving the target data input by the target user; a unit, for creating dialogue information based on the target data and the time tag, and storing the dialogue information in a first storage space; the device also includes: a second determination unit, for determining the output time of the execution result after generating and executing the control instruction and obtaining the execution result; a storage unit, for storing the execution result and the output time in the dialogue information in the first storage space; a screening unit, for screening the output information containing the detection result in the first storage space every second preset time period, and storing the output information of the detection result in the second storage space; a cleaning unit, for cleaning data in the second storage space according to the output time of the execution result every first preset time period.

[0019] Furthermore, the code library contains code information of multiple codes, and the multiple codes include at least one of the following: a first code for controlling the relative movement or absolute movement of the spherical camera; a second code for maintaining the preset point position of the spherical camera; a third code for collecting the status information of the spherical camera, querying the information of the spherical camera and obtaining the shooting results of the spherical camera; a fourth code for performing target detection tasks based on the multiple deep learning models; a fifth code for patrolling the environment in a preset direction or within a preset range; a sixth code for querying information; a seventh code for starting or terminating a task, and the code information includes at least: code function description information, code name, code input parameters, and input parameter information.

[0020] To achieve the above-mentioned objectives, according to one aspect of the present application, a computer program product is provided, including a computer program. When the computer program is executed by a processor, it implements any of the above-mentioned methods for controlling a spherical camera. When the computer program is executed by the processor, it implements the steps of the method for controlling a spherical camera described in each embodiment of the present application.

[0021] To achieve the above objectives, according to one aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes stored computer instructions, wherein when the computer instructions are executed by a processor, any one of the above-mentioned methods for controlling a spherical camera is implemented.

[0022] To achieve the above-mentioned objectives, according to one aspect of the present application, an electronic device is provided, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors implement any of the above-mentioned methods for controlling a spherical camera.

[0023] Through the present application, the following steps are adopted: receiving target data input by a target user, wherein the target data includes at least one of the following: text data, voice data; parsing the pending tasks in the target data through the large language model in the intelligent body model, generating and executing control instructions, and obtaining execution results, wherein the control instructions are instructions for controlling multiple spherical cameras in a camera group based on the codes in the code library, and the code library contains codes for calling multiple deep learning models in the intelligent body model; inputting the execution results into the intelligent body model, updating the control instructions multiple times through the intelligent body model, executing the control instructions after each update, and updating the execution results until the pending tasks are completed, and generating output results, wherein the output results include at least one of the following: text information, image information, and video information, which solves the problem in the related art that when controlling the spherical camera to perform shooting tasks, manual adjustments need to be made repeatedly according to the operating manual and camera model, resulting in low camera working efficiency.

[0024] By receiving and parsing data input by the target user in text or voice format, it can quickly understand the user's specific needs in the video surveillance scenario, achieving the technical effect of efficient task identification. At the same time, the large language model in the intelligent agent model analyzes the target data and automatically generates precise control commands to operate the camera group, allowing the spherical network camera to autonomously complete complex tasks such as target detection and tracking, realizing intelligent monitoring, further achieving the technical effect of improving monitoring efficiency and accuracy, and reducing human intervention. In addition, by continuously feeding back execution results to the intelligent agent model for iterative optimization of control commands, it can refine the understanding of user commands, making the response more closely aligned with the user's actual needs, improving the accuracy and effectiveness of control commands, and enabling the system to adapt to the ever-changing technological environment and user needs, maintaining long-term technological leadership and competitiveness, and improving the system's automation and intelligence levels. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0026] Figure 1 is a flow chart of a method for controlling a spherical camera according to the first embodiment of the present application;

[0027] Figure 2 is a schematic diagram of a processing flow of an optional spherical camera control system provided in accordance with the first embodiment of the present application;

[0028] Figure 3 This is a schematic diagram of an optional execution flow of an intelligent agent provided according to the first embodiment of the present application;

[0029] Figure 4 This is a flow chart of an optional large language model (LLM) provided in accordance with the first embodiment of the present application;

[0030] Figure 5 is a schematic diagram of a control device for a spherical camera according to the second embodiment of the present application;

[0031] Figure 6 Schematic diagram of the control electronic device of the spherical camera provided in Example 5 of the present application. DETAILED DESCRIPTION

[0032] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0033] It should be noted that the user information (including but not limited to user device information, user personal information, collected data, used data, generated data, processed data, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, collected information, used information, generated information, processed information, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data comply with the relevant laws, regulations and standards of relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse. For example, an interface is set up between this system and relevant users or institutions. Before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or institution through the interface, and obtain relevant information after receiving the consent information fed back by the aforementioned user or institution.

[0034] It should be noted that this application provides users with corresponding operation entrances for them to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered.

[0035] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0037] Example 1

[0038] The present invention will be described below in conjunction with preferred implementation steps. Figure 1 Flowchart of the control method of the spherical camera provided in accordance with the first embodiment of the present application. Figure 1 As shown, the method includes the following steps:

[0039] Step S101: receiving target data input by a target user, wherein the target data includes at least one of the following: text data and voice data.

[0040] In this first embodiment, user-provided target data is received as a command source for controlling a dome network camera. The target data can be text or voice data. Users express their monitoring needs through natural language, such as moving the viewing angle, detecting or tracking targets, and other operations. Voice recognition technology is used to convert voice data into text, or the text data is directly processed. A large language model is used to analyze user intent and generate corresponding control logic, enabling intelligent and efficient camera operation to meet the diverse needs of users in video surveillance and intelligent security scenarios. This design simplifies the interaction between users and cameras, improving the user experience and the intelligence of system responses.

[0041] In step S102, the large language model in the intelligent agent model is used to parse the pending tasks in the target data, generate and execute control instructions, and obtain execution results. The control instructions are instructions for controlling multiple spherical cameras in the camera group based on the codes in the code library, and the code library contains codes for calling multiple deep learning models in the intelligent agent model.

[0042] In this first embodiment, the agent model uses a large language model (LLM) to deeply analyze the target data (whether in text or voice) input by the user. This analysis process aims to understand the user's intent and identify the task to be processed, such as adjusting the camera's viewing angle, identifying a specific object, or performing an environmental inspection. The LLM then uses its natural language processing capabilities and industry knowledge base to accurately translate the identified task into computer-executable control instructions.

[0043] Control instructions are generated based on predefined code in a code library. This code library contains code that calls multiple deep learning models within the agent model to perform complex data processing tasks. For example, the object detection model is used to identify and detect unknown objects in video surveillance, while other models, such as image classification and keypoint detection, are used to perform specific data analysis.

[0044] After parsing the control instructions, the intelligent body will call the corresponding code to control the spherical cameras in the camera group to perform the monitoring tasks required by the user, such as adjusting the focal length or pan / tilt position. It may also involve the collaborative work of the entire camera group, such as executing a unified inspection strategy.

[0045] Step S103: input the execution result into the intelligent agent model, update the control instructions multiple times through the intelligent agent model, execute the updated control instructions each time, and update the execution result until the task to be processed is completed, and generate an output result, wherein the output result includes at least one of the following: text information, image information, and video information.

[0046] In this first embodiment, after the agent model executes the control command, it can collect the camera's execution results, including the camera's response status, target detection image data, real-time video monitoring feedback, and other information. These execution results are then re-input into the agent model's LLM for further analysis and decision optimization.

[0047] During multiple rounds of feedback and decision-making iterations, the intelligent agent model dynamically adjusts its control strategy based on the execution results, generating new control instructions. These instructions can, for example, fine-tune the camera's viewing angle, update target detection parameters, perform target tracking, or perform environmental inspections. Each generated control instruction is promptly distributed to the dome cameras in the camera group, enabling more precise monitoring operations. This process ensures that the intelligent agent model can continuously adjust its control logic based on real-time monitoring conditions until the user-defined pending tasks are fully accomplished.

[0048] Ultimately, the agent model generates output in a user-friendly format. This output can include textual information, such as event descriptions or operation status; image information, such as processed screenshots of surveillance footage; or video information, including a complete record of the event and the monitoring process. This multimodal output not only enriches the ways users receive information but also ensures comprehensiveness and intuitiveness, enhancing the system's application value and user experience in fields such as video surveillance and intelligent security.

[0049] In an optional embodiment, the target data input by the user may be: "Turn the field of view to the right, detect whether there is anyone in the field of view, and continue shooting if there is someone". The output results of the intelligent agent model may include: "First thinking, execute the specified operation, execute the move field of view (move view) code, parameters: {"direction": "right", "duration": 3}. Second thinking, execute the specified operation, execute the target detection (object detect) code, parameters: {"object name": "person", "conf (confidence)": 0.5}. Third thinking, execute the specified operation, execute the target detection (Trackobject) code, parameters: {"object name": "person", "conf (confidence)": 0.5}. Fourth thinking, execute the specified operation. Final output: The gimbal has moved to the right for 3 seconds, and a person is detected in the field of view. Now start tracking the person. If you need to change the position of the gimbal or stop tracking, please tell me."

[0050] In summary, the control method of a spherical camera provided in Example 1 of the present application receives target data input by a target user, wherein the target data includes at least one of the following: text data, voice data; parses the pending tasks in the target data through the large language model in the intelligent agent model, generates and executes control instructions, and obtains execution results, wherein the control instructions are instructions for controlling multiple spherical cameras in a camera group based on the codes in the code library, and the code library contains codes for calling multiple deep learning models in the intelligent agent model; inputs the execution results into the intelligent agent model, updates the control instructions multiple times through the intelligent agent model, executes the updated control instructions each time, and updates the execution results until the pending tasks are completed, and generates output results, wherein the output results include at least one of the following: text information, image information, and video information, which solves the problem in the related art that when controlling a spherical camera to perform a shooting task, manual adjustments need to be made repeatedly according to the operating manual and camera model, resulting in low camera working efficiency.

[0051] By receiving and parsing data input by the target user in text or voice format, it can quickly understand the user's specific needs in the video surveillance scenario, achieving the technical effect of efficient task identification. At the same time, the large language model in the intelligent agent model analyzes the target data and automatically generates precise control commands to operate the camera group, allowing the spherical network camera to autonomously complete complex tasks such as target detection and tracking, realizing intelligent monitoring, further achieving the technical effect of improving monitoring efficiency and accuracy, and reducing human intervention. In addition, by continuously feeding back execution results to the intelligent agent model for iterative optimization of control commands, it can refine the understanding of user commands, making the response more closely aligned with the user's actual needs, improving the accuracy and effectiveness of control commands, and enabling the system to adapt to the ever-changing technological environment and user needs, maintaining long-term technological leadership and competitiveness, and improving the system's automation and intelligence levels.

[0052] Optionally, in the control method of the spherical camera provided in Example 1 of the present application, the execution result is input into the intelligent agent model, the control instruction is updated multiple times through the intelligent agent model, the updated control instruction is executed each time, and the execution result is updated until the task to be processed is completed, and the output result is generated, including: inputting the execution result into the large language model for the Nth time, analyzing and making decisions on the target data through the large language model, and obtaining the control instruction after the Nth update; calling the code in the code library according to the control instruction after the Nth update, updating the execution result, and obtaining the Nth execution result corresponding to the control instruction after the Nth update; judging whether to continue to update the control instruction through the large language model based on the task to be processed and the Nth execution result; in the case of updating the control instruction, storing the Nth execution result, and inputting the Nth execution result into the large language model for the N+1th time; in the case of not updating the control instruction, generating the output result based on the execution result corresponding to each execution of the control instruction.

[0053] In this first embodiment, the agent model receives the execution result corresponding to the N-1th control instruction and feeds it into the large language model for the Nth time. It then performs a detailed analysis of the result and intelligently decides on the Nth updated control instruction. Subsequently, based on the Nth updated control instruction, the agent dynamically calls code from the code library, which includes scripts for multiple deep learning models, to control the camera group for more precise operation and obtain the execution result corresponding to the Nth control instruction.

[0054] The agent then determines whether to continue optimizing and updating control instructions based on the degree of match between the pending task and the latest execution result. If an update is necessary, the current execution result is stored and progressively fed back into the large language model for the N+1th decision and control instruction generation, ensuring that each control instruction is the optimal solution based on the latest monitoring conditions.

[0055] Finally, when the intelligent agent determines that there is no need to further update the control instructions, it generates the final output results based on all the execution results corresponding to the previous control instructions. The output results are presented in the form of text, images or videos, and are intuitively fed back to the user to ensure that the user can clearly understand the monitoring progress or abnormal events discovered, realizing the automated closed loop of the monitoring task, and significantly improving the application efficiency and user experience of the system in the fields of video surveillance and intelligent security.

[0056] In an optional embodiment, if the LLM model's analysis results require the execution of related code in the code library, the "Action" and "Action_Input" fields in the instruction can be controlled to parse the code to be called and its input parameters, and the subsequent content can be truncated (because the subsequent content may be an illusion generated by the language model). For example, the following content can be extracted: "Execute move_view code, parameters: {"direction":"right","duration":3}". After the code is executed, the processing results are formatted and converted into the text format required by the LLM model. For example, after the camera rotation code is executed, the output is: "The gimbal has moved to the right for 3 seconds", and the LLM then reconsiders the next disposal strategy.

[0057] If the analysis results of the LLM model do not require the execution of related code in the code library, the model thinking results will be judged. If the output conditions are met: no code execution is required and the "Final_Answer" field exists (equivalent to completing the above-mentioned pending tasks), it will be output directly. If further thinking is required, it will be returned to LLM for processing again.

[0058] Optionally, in the control method of the spherical camera provided in Example 1 of the present application, the large language model in the intelligent agent model is used to parse the tasks to be processed in the target data, generate and execute control instructions, and obtain execution results, including: thinking about the tasks to be processed through the large language model to obtain the output results of the large language model; converting the output results into control instructions according to a preset output format, wherein the control instructions include control operations, and the control instructions include at least the following information: the tasks to be processed, the execution operations, the input parameter information of the execution operations, and the expected execution results of the execution operations; executing the control instructions to determine the execution results.

[0059] In the first embodiment of the present invention, the large language model performs in-depth analysis and intelligent thinking on the task to be processed input by the user to generate a preliminary output result, which includes the understanding of the task and the code control strategy. Subsequently, the intelligent agent model converts the output result of the large language model into an executable control instruction according to the preset output format. The instruction details the specific control operation, including the description of the task to be processed, the action to be performed, the input parameters of the action, and the expected execution result. Among them, the number of control operations contained in the control instruction can be zero, one or more. For example, if the number of control operations in the control instruction is zero, it means that the output result of the large language model is the final output (equivalent to the end of the user conversation); if the number of control operations in the control instruction is multiple, only the first control instruction output can be executed, because the subsequent output of the large language model may be hallucination content.

[0060] In an optional embodiment, the output result of the LLM model may be as follows: "{Question: summarize the question you need to answer, Thought: you should always think about what you should do, Action: the operation to be performed, the operation must be selected from the code provided to you, if the operation does not need to be performed, no output is required, Action_Input: the input of the operation to be performed, if the operation does not require a specified input, {none} is returned, Observation: the return result of the performed operation}, Thought: you know the final answer, Final_Answer: the final result based on the original question". If the user's original question needs to be broken down into multiple steps, "Thought", "Action", "Action_Input" and "Observation" may be output multiple times in the output result, that is, the process of generating "Thought", "Action", "Action_Input" and "Observation" may be executed multiple times or zero times.

[0061] Next, the agent model invokes the camera control code and deep learning model based on the generated control instructions, executes the operations specified in the instructions, monitors the execution process in real time, and collects the execution results, including the specific response of the camera, detected target information, or video analysis results. A control instruction can be expressed as "Action: The operation to be performed. The operation must be selected from the code provided to you. If the operation is not required, no output is required. Action_Input: The input of the operation to be performed. If the operation does not require a specified input, {none} is returned." The control instructions generated by the agent model include at least the following information: the task to be processed, the operation to be performed, the input parameter information for the operation, and the expected execution result of the operation. The structured design of the control instructions ensures that the agent can accurately control the spherical cameras in the camera group and perform tasks such as target detection and viewing angle adjustment.

[0062] Through the above steps, not only the intelligent control of the spherical network camera is achieved, but also the clarity and precision of the control instructions are ensured, which greatly improves the response efficiency of the system and the accuracy of task execution. It provides users with a set of intelligent monitoring solutions with simple operation and instant feedback, further achieving the technical effect of significantly enhancing the application efficiency in the field of video surveillance and intelligent security.

[0063] Optionally, in the control method of a spherical camera provided in the first embodiment of the present application, the large language model in the intelligent agent model is used to parse the tasks to be processed in the target data, generate and execute control instructions, and obtain execution results, including: when the large language model and multiple deep learning models are deployed on a local computing device, the target data is sent to the local computing device, the large language model deployed on the local computing device is used to parse the tasks to be processed in the target data, generate and execute control instructions, and obtain execution results, wherein the computing power of the local computing device is greater than the preset computing power, and the control instructions are used to call the multiple deep learning models in the local computing device to execute tasks; or, when the target model is deployed on a cloud computing device, the large language model is used to parse the tasks to be processed in the target data, generate and execute control instructions, and send the control instructions to the cloud computing device, wherein the target model is a model in which the number of model parameters of the multiple deep learning models is greater than the preset number, and the control instructions are used to call the target model in the cloud computing device to execute tasks; and receive the execution results sent by the cloud computing device.

[0064] In this first embodiment, by flexibly deploying a large language model and multiple deep learning models on local or cloud computing devices, the intelligent agent can effectively process the target data.

[0065] When these models are deployed on a local computing device, if the computing power of the local device exceeds a preset threshold, the target data is sent directly to the local machine. The large language model analyzes the target data for pending tasks, generating control instructions. These instructions are then executed by multiple deep learning models on the local computing device, eliminating network transmission delays and providing real-time results. Local deployment is particularly suitable for scenarios with high response speed requirements, such as emergency monitoring.

[0066] On the other hand, if there is a target model with a large number of model parameters among multiple deep learning models, the target model is placed on a cloud computing device, and the target data is uploaded to the cloud. The local large language model performs in-depth analysis and decision-making on the task to be processed, generates control instructions, and then sends the control instructions to the corresponding target model in the cloud computing device. The cloud has richer computing resources and can support the operation of more complex models, thereby performing tasks with higher precision. After the target model in the cloud computing device executes the control instructions, the intelligent body will receive the execution results returned by the cloud computing device, realizing remote intelligent control. The cloud-based deployment of deep learning models is suitable for environments with limited resources or requiring large-scale data analysis, such as cross-regional monitoring networks.

[0067] In an optional embodiment, the large language model primarily leverages its ability to understand human language, combined with preset text prompts, to decompose the input question into tasks and generate a processing flow. If the large language model is executed on the client (i.e., locally as described above), a relatively small model is selected. If the large language model is executed in the cloud, a model with a larger number of parameters can be selected and then packaged as a service for access by the client-side agent.

[0068] By choosing the above two deployment strategies, based on considerations of computing power requirements and network conditions, the efficient operation and wide applicability of the intelligent control system are ensured, and the intelligence level and user experience in the video surveillance field are improved.

[0069] Optionally, in the control method of the spherical camera provided in Example 1 of the present application, before parsing the task to be processed in the target data through the large language model in the intelligent body model, the above method also includes: processing the target data and the domain knowledge base according to a preset format to obtain prompt words, wherein the prompt words include at least one user intent, so that the large language model parses each user intent in the at least one user intent to obtain the task to be processed; the domain knowledge base at least includes code information corresponding to the control instruction; determining the preset identity information of the intelligent body model based on the target data, and updating the prompt words based on the preset identity information, and inputting the updated prompt words into the large language model.

[0070] In the first embodiment, the agent model first combines the target data with the data in the domain knowledge base and converts it into a preset format to obtain the above-mentioned prompt words. If the user's original question needs to be broken down into multiple steps, the prompt words can include one or more user intents, such as monitoring a specific target, adjusting the shooting angle, or performing abnormal event detection. By breaking down the prompt words, the large language model can deeply analyze each user intent one by one, identify and understand user needs, convert them into tasks to be processed, and prepare to generate corresponding control instructions.

[0071] Then, based on the target data, the agent model determines its pre-set identity information. This identity information is incorporated into the prompt word, guiding the large language model to more accurately simulate the agent's behavioral logic, ensuring that the generated control instructions meet user expectations and industry standards. For example, the identity information might be as follows: "You are an intelligent assistant designed to control spherical network cameras. Your name is 'Smart Pan-Tilt Agent.' You can execute the instructions given by the user. If the user's instructions leave you with incomplete information for decision-making, you can return the information to the user and ask them to provide the necessary information, answering the user's question to the best of your ability."

[0072] Finally, the intelligent agent model inputs the updated prompt words into the large language model, uses the model's deep learning capabilities, combines industry knowledge and user intent, and generates optimized control instructions, thereby realizing intelligent and autonomous control of the spherical network camera, improving the efficiency and accuracy of video surveillance.

[0073] In an optional embodiment, some of the data contained in the domain knowledge base can be represented by the multiple data records shown in Table 1. Here, camera control instructions refer to control instructions for a spherical camera or the multiple deep learning models described above, and code information refers to the code information corresponding to the camera control instructions in the code library. By inputting this code information into the large language model, the large language model can be more accurately instructed to generate clear and precise control instructions.

[0074] Table 1

[0075]

[0076]

[0077] By adding the correspondence between camera control instructions and code information in the code library in the domain knowledge base, an efficient conversion from user instructions to specific execution codes is achieved, ensuring that the intelligent agent can accurately call the corresponding execution script to complete the user-specified operations.

[0078] In an optional embodiment, the large language model can also be tested using historical user conversation data to evaluate the model's performance in a real conversation environment. Based on the test results, the data in the domain knowledge base is dynamically updated to correct the model's deviations in parsing specific instructions and further optimize the model's decision logic. For example, if during debugging, it is discovered that the large language model cannot generate correct thought processes through its own reasoning, a corresponding disposal strategy can be added to the domain knowledge base, using the aforementioned target data and the updated domain knowledge base to generate prompt words, which are then input into the large language model. Through this iterative optimization process, the large language model can continuously evolve to better serve users, achieve intelligent control of spherical network cameras, and improve monitoring efficiency and accuracy.

[0079] Optionally, in the control method of the spherical camera provided in Example 1 of the present application, after receiving the target data input by the target user, the above method further includes: determining the time tag of the target data; creating dialogue information based on the target data and the time tag, and storing the dialogue information in the first storage space; after generating and executing the control instruction to obtain the execution result, the above method further includes: determining the output time of the execution result; storing the execution result and the output time in the dialogue information of the first storage space; filtering the output information containing the detection result in the first storage space every second preset time period, and storing the output information of the detection result in the second storage space; and performing data cleaning in the second storage space according to the output time of the execution result every first preset time period.

[0080] In this first embodiment, after receiving the user's target data, the system automatically determines the data's time stamp to record the precise time of the user's request. The time stamp can be expressed as "[xx year xx month xx day, xx:xx:xx]," thus introducing the concept of time. Subsequently, the target data and time stamp are combined to create conversation information and store it in the system's first storage space, the short-term conversation memory. This short-term conversation memory enhances the agent's ability to retain conversation context.

[0081] After executing a control instruction and obtaining the result, the output time of the result is recorded and stored together with the execution result in the conversation information in the first storage space, forming a complete conversation history. To continuously monitor and analyze anomalies or important events, after a second preset period of time, the output information containing test results in the first storage space is automatically filtered and these key test results are stored in the second storage space, i.e., the long-term event memory.

[0082] For example, when executing a target detection cruise mission, the data can be stored in the long-term event memory according to "time: xxx, camera IP: xxx, camera PTZ value: xx, xx, xx, xxx target detected at position [x1, y1, x2, y2]. Because the long-term event memory includes information such as time annotations and camera angles, it can be compared over a period of time. For example, when the current input is from the last few days or the previous day, the intelligent agent model can match the relevant results based on time and query them. In addition, the long-term memory is cleared according to time (i.e., the first preset duration mentioned above). For example, it is stored for 7*24 hours to remove outdated dialogue information and ensure the real-time and effectiveness of dialogue management.

[0083] The establishment of a long-term event memory library helps the system to track and analyze historical events over a long period of time, providing rich data support for subsequent intelligent decision-making, ensuring the intelligent and autonomous control performance of spherical network cameras in complex monitoring environments, and improving the efficiency and reliability of the overall monitoring system.

[0084] Optionally, in the control method of the spherical camera provided in Example 1 of the present application, the code library contains code information of multiple codes, and the multiple codes include at least one of the following: a first code for controlling the relative movement or absolute movement of the spherical camera; a second code for maintaining the preset position of the spherical camera; a third code for collecting the status information of the spherical camera, querying the information of the spherical camera and obtaining the shooting results of the spherical camera; a fourth code for performing target detection tasks based on multiple deep learning models; a fifth code for patrolling the environment in a preset direction or within a preset range; a sixth code for querying information; a seventh code for starting or terminating a task, and the code information includes at least: code function description information, code name, code input parameters, and input parameter information.

[0085] In the first embodiment of the present invention, the code library includes codes for the intelligent model to call. The code library can also be called a tool set, which includes multiple tools encapsulated into code functions or services. This solution designs the following basic codes for the control scenario of the ball camera, as shown in Table 2, where the code name and code function description are provided as information to the large language model, and the code implementation logic is described in the language. Among them, the first code can include the camera perspective relative movement code, the camera perspective absolute movement code and the lens zoom code, the second code can include setting prefabricated points and setting prefabricated points, the third code is the camera current state acquisition, the fourth code can include target detection code and cloud target analysis code, the fifth code can include inspection code, the sixth code can include camera query code and search engine code, and the seventh code can include camera link code, camera link shutdown code and continuous operation stop code.

[0086] Table 2

[0087]

[0088]

[0089]

[0090] In an optional embodiment, since the use of code needs to be learned through the semantic understanding ability of the large language model, a code specification can be formed in a fixed format, so that when the code needs to be called, the code specification is combined with the control instruction and input into the large language model.

[0091] For example, the following takes the view movement and camera query codes as examples for explanation. The code information is as follows: {'name_for_human':'Camera view movement code', 'name_for_model':'move_view', 'description_for_model':'The camera view movement code can adjust the camera direction by controlling the movement of the camera gimbal, and rotate in the four directions of "left", "right", "up" and "down"', 'parameters': [{'name':'direction', 'description':'The direction of the camera view movement (can be set) "left", "right", "up", "down"), setting it to "left" will control the camera to move left, setting it to "right" will control the camera to move right, setting it to "up" will control the camera to move up, setting it to "down" will control the camera to move down', 'required':True, 'schema':{'type':'string'},}, {'name':'duration', 'description':'Control the movement duration of the gimbal (in seconds, range: 0-7 seconds)', 'required':True, 'schema':{'type':int},}],}.

[0092] Furthermore, the description finally generated by combining the code manual and the control instructions can be as follows: move_view: Calling this code can interact with the camera view movement code API. What is the camera view movement code API used for? The camera view movement code can adjust the camera direction by controlling the movement of the camera gimbal. The code can be used to control the camera to rotate in the four directions of "left", "right", "up" and "down". Input parameters (Parameters): [{'name':'direction','description':'The direction of the camera view movement (can be set to "left", "right", "up", "down"). If set to "left", it will control the camera to move to the left, if set to "right", it will control the camera to move to the right, if set to "up", it will control the camera to move up, and if set to "down", it will control the camera to move down','required':True,'schema':{'type':'string'}},{'name':'duration','description':'Control the movement duration of the gimbal (in seconds, range: 0-7 seconds)','required':True,'schema':{'type':<class'int'>}}] is organized in the format of JSON objects.

[0093] Optionally, in this embodiment 1, Figure 2This is a schematic diagram of the processing flow of an optional dome camera control system according to Example 1 of the present application. Users can choose between voice and text input. Text input can be directly transmitted to the agent for analysis. Voice input requires the ASR (Automatic Speech Recognition) module to convert the user's voice into text before processing. The agent module analyzes the user's input commands and automatically performs operations such as command segmentation, action execution, and session recording. During this module processing, the cameras in the camera group are controlled based on user commands and the results of its own analysis. The camera group is primarily used to network large-scale dome camera equipment, and then the group network is connected to the agent network so that the agent can control a specific camera according to user commands. The agent module requires running a large language model for text analysis, as well as deep learning models in the code base for object detection, key point matching, and other tasks. The camera system requires extremely high response speed for tasks such as event detection and tracking. To reduce system response latency, a certain amount of computing power is configured locally on the agent model. For example, the model used requires at least 30GB of video memory to run. For cost reasons, the local side generally uses low-cost, smaller models for analysis. When local computing resources are limited or a more powerful model is needed for auxiliary analysis, the data will be uploaded to the cloud for collaborative analysis. The LLM module design will encapsulate the tasks that require cloud collaboration into code for system calls. The analysis results of the intelligent agent are divided into three types: text, image, and video stream. The different forms of output are displayed to the user according to the front-end design. The text form will also be converted into speech using the TTS (Text-to-Speech) module, and this speech will then drive the digital human to generate a broadcast video.

[0094] Optionally, in this embodiment 1, Figure 3The following is a schematic diagram of the execution flow of an optional intelligent agent according to the first embodiment of the present application. First, the intelligent agent receives text input from the user, including the user's specific operational requirements or requests for the dome network camera. This text is then managed and stored, added to a historical message database to ensure the continuity of the conversation context and facilitate subsequent analysis and retrieval of historical information. Secondly, based on the user's text input, the intelligent agent integrates the code information corresponding to the camera operation instructions and industry background knowledge in the domain knowledge base, and invokes a large language model (LLM) to deeply process the stored text to improve the accuracy and intelligence of the understanding. The processed results are then stored again in the domain knowledge base, forming a closed-loop learning loop and continuously optimizing the model's decision logic. Finally, the LLM-processed content is parsed to determine whether code invocation is necessary. If the parsing results indicate that code invocation is necessary, the intelligent agent automatically performs the corresponding action, such as adjusting the camera angle or enabling object detection, and sends the feedback results after code execution back to the LLM for a new round of thinking. If code invocation is not necessary, the agent checks to see if the end point of thinking has been reached. If the goal is not reached, dialogue management will continue to deepen understanding and analysis; once the thinking is completed, the intelligent body will output the final conclusion in text form, or convert it into specific operation instructions and send them to the camera to complete the entire intelligent control process.

[0095] Optionally, in this embodiment 1, Figure 4 It is a flow chart of the optional large language model LLM provided according to the first embodiment of the present application. First, the operation rules and strategies are extracted from the camera operation thinking knowledge base, including the conventional control logic of the spherical network camera, such as preset point management, viewing angle adjustment, etc., to provide a basic control thinking framework for the intelligent agent. Then, based on the above knowledge, the complex control logic of the camera is determined and set, including target detection logic, inspection logic and target tracking strategy, etc., to cope with more complex monitoring needs and realize the advanced control of the camera by the intelligent agent. Secondly, the integrated thinking logic settings, intelligent agent personality information and complex camera control logic are input into the large language model (LLM_base), that is, the execution logic and role settings of specific tasks are set, so that the model can more accurately understand and generate control instructions that meet the characteristics of the intelligent agent. Finally, the LLM_base module receives the content output by the code library (such as pan / tilt control status, target detection results, etc.), performs in-depth processing based on the input information, and generates responses to user input or control instructions for the camera. This enables the entire system to intelligently analyze user needs and autonomously call functions in the code library, such as performing open-set target detection and querying weather information, thereby realizing intelligent control and efficient monitoring of spherical network cameras.

[0096] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0097] Example 2

[0098] The second embodiment of the present application further provides a control device for a spherical camera. It should be noted that the control device for a spherical camera of the second embodiment of the present application can be used to execute the control method for a spherical camera provided in the first embodiment of the present application. The control device for a spherical camera provided in the second embodiment of the present application is described below.

[0099] Figure 5 Schematic diagram of the control device of the spherical camera provided in accordance with the second embodiment of the present application. Figure 5 As shown, the device includes: a receiving unit 501, a first executing unit 502 and a second executing unit 503.

[0100] Specifically, the receiving unit 501 is configured to receive target data input by a target user, wherein the target data includes at least one of the following: text data and voice data.

[0101] The first execution unit 502 is used to parse the pending tasks in the target data through the large language model in the intelligent agent model, generate and execute control instructions, and obtain execution results, wherein the control instructions are instructions for controlling multiple spherical cameras in the camera group according to the codes in the code library, and the code library contains codes for calling multiple deep learning models in the intelligent agent model.

[0102] The second execution unit 503 is used to input the execution result into the intelligent agent model, update the control instructions multiple times through the intelligent agent model, execute the control instructions after each update, and update the execution result until the task to be processed is completed, and generate an output result, wherein the output result includes at least one of the following: text information, image information, and video information.

[0103] The control device for a spherical camera provided in the second embodiment of the present application receives target data input by a target user through a receiving unit 501, wherein the target data includes at least one of the following: text data, voice data; the first execution unit 502 parses the task to be processed in the target data through the large language model in the intelligent agent model, generates and executes control instructions, and obtains an execution result, wherein the control instruction is an instruction to control multiple spherical cameras in the camera group according to the code in the code library, and the code library contains codes for calling multiple deep learning models in the intelligent agent model; the second execution unit 503 inputs the execution result into the intelligent agent model, updates the control instruction multiple times through the intelligent agent model, executes the control instruction after each update, and updates the execution result until the task to be processed is completed, and generates an output result, wherein the output result includes at least one of the following: text information, image information, and video information, which solves the problem in the related art that when controlling the spherical camera to perform shooting tasks, manual adjustments need to be made repeatedly according to the operating manual and camera model, resulting in low camera working efficiency.

[0104] By receiving and parsing data input by the target user in text or voice format, it can quickly understand the user's specific needs in the video surveillance scenario, achieving the technical effect of efficient task identification. At the same time, the large language model in the intelligent agent model analyzes the target data and automatically generates precise control commands to operate the camera group, allowing the spherical network camera to autonomously complete complex tasks such as target detection and tracking, realizing intelligent monitoring, further achieving the technical effect of improving monitoring efficiency and accuracy, and reducing human intervention. In addition, by continuously feeding back execution results to the intelligent agent model for iterative optimization of control commands, it can refine the understanding of user commands, making the response more closely aligned with the user's actual needs, improving the accuracy and effectiveness of control commands, and enabling the system to adapt to the ever-changing technological environment and user needs, maintaining long-term technological leadership and competitiveness, and improving the system's automation and intelligence levels.

[0105] Optionally, in the control device of the spherical camera provided in the second embodiment of the present application, the above-mentioned second execution unit 503 includes: a first generation subunit, which is used to input the execution result into the large language model for the Nth time, analyze and make decisions on the target data through the large language model, and obtain the control instruction after the Nth update; an updating subunit, which is used to call the code in the code library according to the control instruction after the Nth update, update the execution result, and obtain the Nth execution result corresponding to the control instruction after the Nth update; a judgment subunit, which is used to judge whether to continue to update the control instruction through the large language model based on the task to be processed and the Nth execution result; a storage subunit, which is used to store the Nth execution result when the control instruction is updated, and input the Nth execution result into the large language model for the N+1th time; and a second generation subunit, which is used to generate an output result based on the execution result corresponding to each execution of the control instruction without updating the control instruction.

[0106] Optionally, in the control device of the spherical camera provided in Example 2 of the present application, the above-mentioned first execution unit 502 includes: a third generation subunit, used to think about the task to be processed through a large language model to obtain the output result of the large language model; a conversion subunit, used to convert the output result into a control instruction according to a preset output format, wherein the control instruction includes a control operation, and the control instruction includes at least the following information: the task to be processed, the execution operation, the input parameter information of the execution operation, and the expected execution result of the execution operation; a determination subunit, used to execute the control instruction and determine the execution result.

[0107] Optionally, in the control device of the spherical camera provided in Example 2 of the present application, the above-mentioned first execution unit includes: a first execution sub-unit, which is used to send the target data to the local computing device when the large language model and multiple deep learning models are deployed on the local computing device, and parse the tasks to be processed in the target data through the large language model deployed in the local computing device, generate and execute control instructions, and obtain execution results, wherein the computing power of the local computing device is greater than the preset computing power, and the control instructions are used to call the multiple deep learning models in the local computing device to execute tasks; or, a second execution sub-unit, which is used to parse the tasks to be processed in the target data through the large language model when the target model is deployed on a cloud computing device, generate and execute control instructions, and send the control instructions to the cloud computing device, wherein the target model is a model in which the number of model parameters of the multiple deep learning models is greater than the preset number, and the control instructions are used to call the target model in the cloud computing device to execute tasks; and receive the execution results sent by the cloud computing device.

[0108] Optionally, in the control device of the spherical camera provided in Example 2 of the present application, the above-mentioned device also includes: a processing unit, used to process the target data and the domain knowledge base according to a preset format to obtain a prompt word, wherein the prompt word includes at least one user intent, so that the large language model parses each user intent in the at least one user intent to obtain a task to be processed; the domain knowledge base at least includes code information corresponding to the control instruction; an input unit, used to determine the preset identity information of the intelligent body model based on the target data, and update the prompt word based on the preset identity information, and input the updated prompt word into the large language model.

[0109] Optionally, in the control device of the spherical camera provided in Example 2 of the present application, the above-mentioned device also includes: a first determination unit, for determining the time tag of the target data after receiving the target data input by the target user; a unit, for creating dialogue information based on the target data and the time tag, and storing the dialogue information in the first storage space; the device also includes: a second determination unit, for determining the output time of the execution result after generating and executing the control instruction and obtaining the execution result; a storage unit, for storing the execution result and the output time in the dialogue information of the first storage space; a screening unit, for screening the output information containing the detection result in the first storage space every second preset time period, and storing the output information of the detection result in the second storage space; a cleaning unit, for cleaning data in the second storage space according to the output time of the execution result every first preset time period.

[0110] Optionally, in the control device of the spherical camera provided in Example 2 of the present application, the above-mentioned code library contains code information of multiple codes, and the multiple codes include at least one of the following: a first code for controlling the relative movement or absolute movement of the spherical camera; a second code for maintaining the preset position of the spherical camera; a third code for collecting the status information of the spherical camera, querying the information of the spherical camera and obtaining the shooting results of the spherical camera; a fourth code for performing target detection tasks based on multiple deep learning models; a fifth code for patrolling the environment in a preset direction or within a preset range; a sixth code for querying information; a seventh code for starting or terminating a task, and the code information includes at least: code function description information, code name, code input parameters, and input parameter information.

[0111] The dome camera control device includes a processor and memory. The aforementioned receiving unit 501, first execution unit 502, and second execution unit 503 are stored in the memory as program units. The processor executes these program units stored in the memory to implement the corresponding functions. The processor includes a kernel, which retrieves the corresponding program units from the memory. One or more kernels can be provided, and kernel parameters can be adjusted to improve camera operating efficiency. The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory includes at least one memory chip.

[0112] A third embodiment of the present invention provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, a method for controlling a spherical camera is implemented.

[0113] A fourth embodiment of the present invention provides a processor, which is used to run a program, wherein the program executes a method for controlling a spherical camera when the program is running.

[0114] Figure 6 Schematic diagram of the electronic device for controlling a spherical camera according to the fifth embodiment of the present application. Figure 6 As shown, embodiment five of the present invention provides an electronic device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, the following steps are implemented: receiving target data input by a target user, wherein the target data includes at least one of the following: text data, voice data; parsing the task to be processed in the target data through the large language model in the intelligent agent model, generating and executing control instructions, and obtaining an execution result, wherein the control instruction is an instruction for controlling multiple spherical cameras in a camera group according to the code in the code library, and the code library contains codes for calling multiple deep learning models in the intelligent agent model; inputting the execution result into the intelligent agent model, updating the control instruction multiple times through the intelligent agent model, executing the control instruction after each update, and updating the execution result until the task to be processed is completed, and generating an output result, wherein the output result includes at least one of the following: text information, image information, and video information.

[0115] When the processor executes the program, it also implements the following steps: inputting the execution result into the intelligent agent model, updating the control instruction multiple times through the intelligent agent model, executing the control instruction after each update, and updating the execution result until the pending task is completed, and generating an output result, including: inputting the execution result into the large language model for the Nth time, analyzing and making decisions on the target data through the large language model, and obtaining the control instruction after the Nth update; calling the code in the code library according to the control instruction after the Nth update, updating the execution result, and obtaining the Nth execution result corresponding to the control instruction after the Nth update; judging whether to continue to update the control instruction through the large language model based on the pending task and the Nth execution result; in the case of updating the control instruction, storing the Nth execution result, and inputting the Nth execution result into the large language model for the N+1th time; in the case of not updating the control instruction, generating an output result based on the execution result corresponding to each execution of the control instruction.

[0116] When the processor executes the program, it also implements the following steps: parsing the tasks to be processed in the target data through the large language model in the intelligent agent model, generating and executing control instructions, and obtaining execution results, including: thinking about the tasks to be processed through the large language model to obtain the output results of the large language model; converting the output results into control instructions according to a preset output format, wherein the control instructions include control operations, and the control instructions include at least the following information: the tasks to be processed, the execution operations, the input parameter information of the execution operations, and the expected execution results of the execution operations; executing the control instructions to determine the execution results.

[0117] When the processor executes the program, the following steps are also implemented: parsing the pending tasks in the target data through the large language model in the intelligent agent model, generating and executing control instructions, and obtaining execution results, including: when the large language model and multiple deep learning models are deployed on the local computing device, sending the target data to the local computing device, parsing the pending tasks in the target data through the large language model deployed in the local computing device, generating and executing control instructions, and obtaining execution results, wherein the computing power of the local computing device is greater than the preset computing power, and the control instructions are used to call the multiple deep learning models in the local computing device to execute tasks; when the target model is deployed on a cloud computing device, parsing the pending tasks in the target data through the large language model, generating and executing control instructions, and sending the control instructions to the cloud computing device, wherein the target model is a model in which the number of model parameters in the multiple deep learning models is greater than the preset number, and the control instructions are used to call the target model in the cloud computing device to execute tasks; receiving the execution results sent by the cloud computing device.

[0118] When the processor executes the program, the following steps are also implemented: before parsing the pending tasks in the target data through the large language model in the intelligent agent model, the above method also includes: processing the target data and the domain knowledge base according to a preset format to obtain prompt words, wherein the prompt words include at least one user intent, so that the large language model parses each user intent in the at least one user intent to obtain the pending tasks; the domain knowledge base at least includes code information corresponding to the control instruction; determining the preset identity information of the intelligent agent model based on the target data, updating the prompt words based on the preset identity information, and inputting the updated prompt words into the large language model.

[0119] When the processor executes the program, the following steps are also implemented: after receiving the target data input by the target user, the above method also includes: determining the time tag of the target data; creating dialogue information based on the target data and the time tag, and storing the dialogue information in the first storage space; after generating and executing the control instruction to obtain the execution result, the above method also includes: determining the output time of the execution result; storing the execution result and the output time in the dialogue information in the first storage space; filtering the output information containing the detection result in the first storage space every second preset time period, and storing the output information of the detection result in the second storage space; and performing data cleaning in the second storage space based on the output time of the execution result every first preset time period.

[0120] When the processor executes the program, the following steps are also implemented: the code library contains code information of multiple codes, and the multiple codes include at least one of the following: a first code for controlling the relative movement or absolute movement of the spherical camera; a second code for maintaining the preset point position of the spherical camera; a third code for collecting the status information of the spherical camera, querying the information of the spherical camera, and obtaining the shooting results of the spherical camera; a fourth code for performing target detection tasks based on multiple deep learning models; a fifth code for inspecting the environment in a preset direction or within a preset range; a sixth code for querying information; a seventh code for starting or terminating a task, and the code information includes at least: code function description information, code name, code input parameters, and input parameter information.

[0121] The devices in this article can be servers, PCs, PADs, mobile phones, etc.

[0122] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0123] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function specified in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps of the functions specified in the block or blocks. In a typical configuration, the computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory. The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0124] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves. It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0125] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0126] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A method for controlling a spherical camera, characterized in that: include: receiving target data input by a target user, wherein the target data includes at least one of the following: text data and voice data; parsing the pending tasks in the target data using a large language model in the agent model, generating and executing control instructions to obtain execution results, wherein the control instructions are instructions for controlling multiple spherical cameras in the camera group based on code in a code library, the code library containing code for calling multiple deep learning models in the agent model; The execution result is input into the intelligent agent model, the control instruction is updated multiple times through the intelligent agent model, the control instruction after each update is executed, and the execution result is updated until the pending task is completed, and an output result is generated, wherein the output result includes at least one of the following: text information, image information, and video information.

2. The method according to claim 1, characterized in that Inputting the execution result into the agent model, updating the control instruction multiple times through the agent model, and executing the updated control instruction each time, The execution result is updated until the pending task is completed, and the output result is generated, including: Inputting the execution result into the large language model for the Nth time, analyzing and making decisions based on the target data by the large language model, and obtaining the control instruction after the Nth update; Calling the code in the code library according to the control instruction after the Nth update, updating the execution result, and obtaining the Nth execution result corresponding to the control instruction after the Nth update; determining whether to continue updating the control instruction using the large language model according to the task to be processed and the Nth execution result; In the case of updating the control instruction, storing the Nth execution result, and inputting the Nth execution result into the large language model for the N+1th time; In the case of not updating the control instruction, the output result is generated according to the execution result corresponding to each execution of the control instruction.

3. The method according to claim 1, characterized in that The large language model in the agent model is used to parse the tasks to be processed in the target data, generate and execute control instructions, and obtain execution results, including: Considering the task to be processed through the large language model to obtain an output result of the large language model; Converting the output result into the control instruction according to a preset output format, wherein the control instruction includes a control operation, and the control instruction includes at least the following information: the task to be processed, the execution operation, input parameter information of the execution operation, and the expected execution result of the execution operation; Execute the control instruction and determine the execution result.

4. The method according to claim 1, wherein The large language model in the agent model is used to parse the tasks to be processed in the target data, generate and execute control instructions, and obtain execution results, including: In a case where the large language model and the multiple deep learning models are deployed on a local computing device, the target data is sent to the local computing device, the tasks to be processed in the target data are parsed by the large language model deployed in the local computing device, the control instruction is generated and executed, and the execution result is obtained, wherein the computing power of the local computing device is greater than the preset computing power, and the control instruction is used to call the multiple deep learning models in the local computing device to execute the task; or When the target model is deployed on a cloud computing device, the task to be processed in the target data is parsed by the large language model, the control instruction is generated and executed, and the control instruction is sent to the cloud computing device, wherein the target model is a model in which the number of model parameters among the multiple deep learning models is greater than a preset number, and the control instruction is used to call the target model in the cloud computing device to execute the task; and receive the execution result sent by the cloud computing device.

5. The method according to claim 1, wherein Before parsing the task to be processed in the target data using the large language model in the agent model, the method further includes: Processing the target data and the domain knowledge base according to a preset format to obtain prompt words, wherein the prompt words include at least one user intent, so that the large language model parses each user intent in the at least one user intent to obtain the task to be processed; the domain knowledge base includes at least code information corresponding to the control instruction; The preset identity information of the agent model is determined according to the target data, the prompt word is updated according to the preset identity information, and the updated prompt word is input into the large language model.

6. The method according to claim 1, characterized in that After receiving target data input by the target user, the method further includes: determining a time tag of the target data; creating conversation information based on the target data and the time tag, and storing the conversation information in a first storage space; After generating and executing the control instruction and obtaining the execution result, the method further includes: determining the output time of the execution result; storing the execution result and the output time in the conversation information of the first storage space; filtering the output information containing the detection result in the first storage space every second preset time period, and storing the output information of the detection result in the second storage space; and cleaning data in the second storage space according to the output time of the execution result every first preset time period.

7. The method according to claim 1, characterized in that The code library includes code information of a plurality of codes, wherein the plurality of codes include at least one of the following: a first code for controlling the relative movement or absolute movement of the spherical camera; a second code for maintaining a preset point position of the spherical camera; a third code for collecting status information of the spherical camera, querying information of the spherical camera, and obtaining shooting results of the spherical camera; and a fourth code for performing a target detection task based on the multiple deep learning models; The fifth code is used to inspect the environment in a preset direction or within a preset range; The sixth code is used to query information; The seventh code is used to start or terminate a task, and the code information includes at least: code function description information, code name, code input parameters, and input parameter information.

8. A control device for a spherical camera, characterized in that: include: a receiving unit, configured to receive target data input by a target user, wherein the target data includes at least one of the following: text data, voice data; a first execution unit, configured to parse the to-be-processed tasks in the target data using a large language model in the agent model, generate and execute control instructions, and obtain execution results, wherein the control instructions are instructions for controlling a plurality of spherical cameras in a camera group according to code in a code library containing code for calling a plurality of deep learning models in the agent model; A second execution unit is used to input the execution result into the intelligent agent model, update the control instruction multiple times through the intelligent agent model, execute the control instruction after each update, and update the execution result until the task to be processed is completed, and generate an output result, wherein the output result includes at least one of the following: text information, image information, and video information.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes stored computer instructions, wherein when the computer instructions are executed by one or more processors, the method for controlling a spherical camera according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: The device comprises one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the control method of the spherical camera according to any one of claims 1 to 7.