Visual task processing method and visual task processing system applied to intelligent terminal

CN122530772APending Publication Date: 2026-08-07ALIBABA CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]然而,上述在进行视觉任务处理时,先接收语音任务再基于语音任务采集图像进行视觉任务处理,导致采集的图像和当时的语音任务不匹配,图像采集存在滞后性,视觉任务处理的实时性较差,且准确率较差

Benefits of technology

[0027]本说明书实施例提供了一种应用于智能终端的视觉任务处理方法,获取目标视觉任务的任务参数,其中,所述任务参数包括所述目标视觉任务的任务信息和对应的参考时间范围,所述参考时间范围为所述目标视觉任务之前的时间范围;根据所述参考时间范围从参考图像集中筛选出目标图像序列,其中,所述参考图像集中存储有所述目标视觉任务之前按时间顺序采集到的至少一帧图像,所述目标图像序列通过所述智能终端上与所述目标视觉任务匹配的摄像头采集;根据所述目标图像序列和所述任务信息,通过目标视觉处理模型,获得所述目标视觉任务的目标处理结果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122530772A_ABST
    Figure CN122530772A_ABST
Patent Text Reader

Abstract

Embodiments of the present specification provide a visual task processing method and a visual task processing system applied to a smart terminal, wherein the visual task processing method comprises: obtaining task parameters of a target visual task, wherein the task parameters comprise task information of the target visual task and a corresponding reference time range, and the reference time range is a time range before the target visual task; selecting a target image sequence from a reference image set according to the reference time range, wherein the reference image set stores at least one image frame collected in time sequence before the target visual task; and obtaining a target processing result of the target visual task by a target visual processing model according to the target image sequence and the task information. When performing visual task processing, the target image sequence corresponding to the reference time range can be selected from the image frames collected before the target visual task as reference information for processing the visual task, thereby ensuring the real-time performance, processing efficiency and accuracy of the visual task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of artificial intelligence technology, and in particular to a visual task processing method and system for use in smart terminals. Background Technology

[0002] With the rapid development of computer technology and artificial intelligence technology, deep learning models have shown great potential in visual task processing, and more and more smart terminals are deploying visual task processing systems based on visual processing models.

[0003] In existing technologies, users typically first issue voice tasks to the smart terminal. Upon receiving the voice task, the smart terminal then controls the camera to capture images. In conjunction with the voice task, the smart terminal can analyze the currently captured images through a visual processing model, providing visual task processing capabilities.

[0004] However, the aforementioned approach to visual task processing, which involves receiving the speech task first and then acquiring images based on it for visual task processing, results in a mismatch between the acquired images and the speech task at that time. Image acquisition is delayed, leading to poor real-time performance and accuracy in visual task processing. Therefore, there is an urgent need for a visual task processing solution that can improve both real-time performance and accuracy. Summary of the Invention

[0005] In view of this, embodiments of this specification provide a visual task processing method applied to a smart terminal. One or more embodiments of this specification also relate to a visual task processing system, an information processing method based on a visual processing model, a task platform, a computing device, a smart terminal, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.

[0006] According to a first aspect of the embodiments of this specification, a visual task processing method applied to a smart terminal is provided, comprising:

[0007] Obtain task parameters for the target visual task, wherein the task parameters include task information of the target visual task and a corresponding reference time range, and the reference time range is the time range prior to the target visual task;

[0008] The target image sequence is selected from the reference image set according to the reference time range, wherein the reference image set stores at least one frame of image acquired in chronological order before the target visual task, and the target image sequence is acquired by a camera on the smart terminal that matches the target visual task;

[0009] Based on the target image sequence and the task information, the target processing result of the target visual task is obtained through the target visual processing model.

[0010] According to a second aspect of the embodiments of this specification, a visual task processing system is provided, deployed on a smart terminal, comprising:

[0011] The task receiving module is configured to acquire task parameters of a target visual task, wherein the task parameters include task information of the target visual task and a corresponding reference time range, and the reference time range is the time range prior to the target visual task.

[0012] An image filtering module is configured to filter a target image sequence from a reference image set according to the reference time range, wherein the reference image set stores at least one frame of image acquired in chronological order prior to the target visual task, and the target image sequence is acquired by a camera on the smart terminal that matches the target visual task;

[0013] The visual reasoning module is configured to obtain the target processing result of the target visual task through a target visual processing model based on the target image sequence and the task information.

[0014] According to a third aspect of the embodiments of this specification, an information processing method based on a visual processing model is provided, applied to a task platform, comprising:

[0015] Receive a model request sent by a smart terminal, wherein the model request includes at least one of the following: scene identifier of the target scene, scene input data of the target scene, and model specification parameters;

[0016] Based on the model request, a corresponding target visual processing model is determined from at least one visual processing model, wherein the target visual processing model is used to execute the above-described visual task processing method applied to a smart terminal.

[0017] According to a fourth aspect of the embodiments of this specification, a computing device is provided, comprising:

[0018] Memory and processor;

[0019] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, they implement the steps of the above-mentioned visual task processing method or information processing method based on a visual processing model applied to a smart terminal.

[0020] According to a fifth aspect of the embodiments of this specification, a smart terminal with a camera is provided, the smart terminal further comprising a voice acquisition device, a visual task processing system and a player;

[0021] The camera is configured to capture at least one frame of image and write it into a reference image set;

[0022] The voice acquisition device is configured to acquire task voice and generate target visual task;

[0023] The visual task processing system is configured to acquire task parameters of the target visual task, wherein the task parameters include task information of the target visual task and a corresponding reference time range, the reference time range being the time range preceding the target visual task; select a target image sequence from a reference image set according to the reference time range; and obtain the target processing result of the target visual task through a target visual processing model based on the target image sequence and the task information.

[0024] The player is configured to play the target processing results of the target visual task.

[0025] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions, which, when executed by a processor, implement the steps of the above-described visual task processing method applied to a smart terminal or the information processing method based on a visual processing model.

[0026] According to a seventh aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described visual task processing method applied to a smart terminal or the information processing method based on a visual processing model.

[0027] This specification provides a visual task processing method for a smart terminal, comprising: acquiring task parameters of a target visual task, wherein the task parameters include task information of the target visual task and a corresponding reference time range, the reference time range being the time range preceding the target visual task; filtering a target image sequence from a reference image set according to the reference time range, wherein the reference image set stores at least one frame of image acquired in chronological order preceding the target visual task, the target image sequence being acquired by a camera on the smart terminal matching the target visual task; and obtaining the target processing result of the target visual task through a target visual processing model based on the target image sequence and the task information.

[0028] One embodiment of this specification implements a method to acquire task information of a target visual task and a reference time range prior to the target visual task. Then, based on the reference time range, a corresponding target image sequence is selected from a reference image set. A target visual processing model is used to perform inference analysis on the selected target image sequence and task information to obtain the corresponding target processing result. The reference image set stores at least one frame of image acquired chronologically prior to the target visual task. Thus, during visual task processing, a target image sequence corresponding to the reference time range can be selected from multiple frames acquired prior to the target visual task as reference information for processing the visual task. This avoids image acquisition lag, ensures the real-time performance of the visual task, guarantees processing efficiency, and improves the accuracy of visual task processing by enabling task processing based on the target image sequence corresponding to the reference time range. Attached Figure Description

[0029] Figure 1 This is a schematic diagram of an image-based visual question-answering process provided in one embodiment of this specification;

[0030] Figure 2 This is a schematic diagram of a video-based visual question-answering process provided in one embodiment of this specification;

[0031] Figure 3 This is an application architecture diagram of a vision task provided by one embodiment of this specification;

[0032] Figure 4 This is a flowchart of a visual task processing method applied to a smart terminal, provided in one embodiment of this specification.

[0033] Figure 5 This is a schematic diagram of the processing procedure of a vision task processing system provided in one embodiment of this specification;

[0034] Figure 6 This is a flowchart illustrating the processing procedure of a visual task processing method applied to a smart terminal, as provided in one embodiment of this specification.

[0035] Figure 7 This is a flowchart illustrating an information processing method based on a visual processing model, provided in one embodiment of this specification.

[0036] Figure 8 This is a schematic diagram of the structure of a task platform provided in one embodiment of this specification;

[0037] Figure 9 This is a schematic diagram of the structure of a vision task processing system provided in one embodiment of this specification;

[0038] Figure 10This is a structural block diagram of a computing device provided in one embodiment of this specification;

[0039] Figure 11 This is a structural block diagram of a smart terminal with a camera provided in one embodiment of this specification. Detailed Implementation

[0040] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0041] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0042] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0043] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0044] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.

[0045] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as natural language processing tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0046] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0047] Large-scale visual models refer to deep learning models designed in the field of computer vision to process and understand images, videos, and text content. These models typically have a large number of parameters, can be trained on large-scale datasets, and can perform a variety of complex visual tasks, such as visual question answering, object recognition, scene understanding, image generation, image description generation (i.e., generating corresponding text descriptions given an image), object detection, and semantic segmentation.

[0048] Visual Question Answering (VQA) is a multimodal artificial intelligence application that combines computer vision and natural language processing. It allows users to ask questions about images or videos and receive corresponding answers. The core of VQA lies in understanding the relationship between the visual content in an image or video and the question, and how to synthesize these two types of information to generate the correct answer.

[0049] Visual Question Answering (VQA) systems are systems that use VQA technology to answer questions based on images or videos.

[0050] Image-based Visual Quality Assurance (VQA): Input an image into a large visual model, and use the large visual model to analyze the image and output the corresponding answer.

[0051] Video VQA: Input a video stream into a large visual model, and use the large visual model to analyze the video stream and output the corresponding answers.

[0052] Real-time performance of visual question-answering systems: the time taken from when a user finishes asking a question to when they receive the corresponding answer.

[0053] The 3A of sound refers to three key technologies in audio processing: Automatic Gain Control (AGC), Automatic Noise Suppression (ANS), and Automatic Echo Cancellation (AEC). Automatic Gain Control (AGC) adjusts the volume level of the audio signal to ensure the output sound remains within a comfortable audible range. It helps resolve issues caused by variations in speaker distance from the microphone or changes in ambient volume. Automatic Noise Suppression (ANS) aims to identify and reduce background noise, such as fan noise, keyboard clicks, or other non-speech sounds, thereby improving speech intelligibility. Echo Cancellation (AEC) removes echoes caused by sound played through speakers being picked up again by the microphone.

[0054] Voice Activity Detection (VAD) is a technique used to distinguish between speech and non-speech elements (such as silence or background noise). It is widely used in various speech processing systems, such as teleconferencing, speech recognition, speech coding, and real-time communication applications. The main purpose of VAD is to accurately detect the actual time periods during which someone is speaking in an audio stream, thereby enabling more efficient processing of audio data, reducing unnecessary computational resource consumption, and improving the efficiency and accuracy of subsequent processing steps. Voice VAD can detect the start and end points of speech.

[0055] Image features: In computer vision, image features refer to attributes or patterns that describe specific information in an image. These features can be low-level (such as edges, corners, color histograms) or high-level (such as objects, scenes, textures). Image feature extraction is fundamental to many computer vision tasks, including image classification, object detection, image retrieval, and image segmentation.

[0056] Image sharpness features: These are image features used to describe whether an image is sharp or not. The most common feature is the image edge. Generally speaking, an image with sharp edges is considered to have higher sharpness.

[0057] Image sharpness detection algorithms: Common ones include gradient-based detection methods, such as the Sobel operator and Canny edge detection; and frequency-based detection algorithms, such as Fourier transform and DCT (discrete cosine transform).

[0058] Image complexity features: These are image features used to describe the complexity of an image. Image complexity can be measured based on various features, including but not limited to texture, color diversity, edge density, spatial frequency distribution, and semantic content. Image edges are commonly used; generally, the more edges and details an image has, the more complex it is considered. Image complexity can be assessed by detecting the number and intensity of edges in the image.

[0059] Image complexity detection algorithms: Image complexity detection is an important task in the field of computer vision. It aims to quantify the image complexity features such as the structural and content complexity of an image. Specifically, it can also adopt gradient-based detection methods, such as the Sobel operator and Canny edge detection; frequency-based detection algorithms, such as Fourier transform and DCT (discrete cosine transform).

[0060] Attitude sensor: A device used to detect and measure the position, orientation, and rotation of an object in three-dimensional space. It is an inertial sensor that typically combines multiple types of sensing elements, such as accelerometers and gyroscopes. It is commonly used to detect the attitude and motion of objects and is widely used in smart terminals such as AR (Augmented Reality) devices and wearable devices to provide more accurate and reliable data.

[0061] Image similarity features can be divided into two aspects: global features and local features. Global features refer to the feature extraction of the entire image, commonly including color histograms and color moments. Local features include SIFT (Scale Invariant Feature Transform) features and SURF (Speed-Up Robust Feature Transform) features.

[0062] Image similarity detection algorithms: Feature-based algorithms commonly used include color histogram, gray-level co-occurrence matrix (GLCM), SIFT (scale invariant feature transform), etc.

[0063] Automatic Speech Recognition (ASR) is a technology that converts speech into text. It combines multiple technologies such as signal processing, pattern recognition, and machine learning, and is widely used in various fields such as smart assistants, voice input methods, telephone customer service systems, and meeting recording.

[0064] Text-to-Speech (TTS) is a technology that converts text into speech, enabling computers and devices to "speak," thus making human-computer interaction more user-friendly. TTS systems are widely used in intelligent assistants, navigation systems, audiobooks, assistive technologies, and other fields, providing convenient services to various users.

[0065] Temporal redundancy in video: The image features of consecutive frames in a video usually change very little, resulting in a large amount of similarity or repetitive information between adjacent frames in a video sequence. This redundancy mainly comes from the fact that objects and scenes do not change much in a short period of time, resulting in a high correlation between consecutive image frames. Understanding and utilizing temporal redundancy is crucial for tasks such as video compression, encoding, processing, and analysis, as it can help reduce data volume, improve transmission efficiency, and optimize storage space.

[0066] Real-time computing power of the model: The video stream that can be processed in real time, expressed as resolution@frame rate. For example, 1080P@5fps means that the model can process a video stream with a resolution of 1080P and a frame rate of 5fps in real time.

[0067] SLLM: Large model with small parameters, typically with a parameter count of less than 10 bytes.

[0068] It should be noted that with the development of large-scale model technology, more and more smart terminals (such as AI glasses, AIRabbit, and AI wearable devices) have incorporated visual question answering systems based on visual large-scale models. By capturing images through the camera on the smart terminal and combining them with voice input questions, the smart terminal can provide visual question answering capabilities through visual large-scale models.

[0069] In one implementation, the visual question answering system can support image-based VQA. Figure 1 This is a schematic diagram of an image-based visual question-answering process provided in one embodiment of this specification, such as... Figure 1 As shown, for the image-based VQA solution, the microphone collects user-side audio in real time and transmits the audio stream to the Automatic Speech Recognition (ASR) module of the visual question-answering system. The ASR module automatically converts the audio into text, which is then transmitted to the intent recognition module. If the intent recognition module determines that the intent is a VQA intent, it issues a command prompt (CMD) to control the camera to take a picture. The image captured by the camera is uploaded to the intent recognition module, which then sends the text and image to the Visual Large Model (VL model) for text question analysis and outputs the answer text. The text-to-speech (TTS) module then converts the answer text into the answer speech and returns it to the player for playback.

[0070] As can be seen from the above, image VQA suffers from image lag (taking a photo must be done after a voice command), poor real-time performance (it requires intent analysis, then taking a photo, transmitting the image, processing it through the model, and then returning the result), inability to provide image context (only the current photo is available), and the ability to ask questions about the current scene (it cannot ask questions about past times). It also suffers from high latency, poor accuracy, and poor flexibility.

[0071] In another implementation, video VQA can solve the above problems of image VQA. The continuous video stream ensures real-time performance (video is transmitted in real time, and the video frame has arrived at the large model when the user asks a question) and context, thereby ensuring the accuracy of visual question answering. Figure 2 This is a schematic diagram of a video-based visual question-answering process provided in one embodiment of this specification, such as... Figure 2 As shown, for the video VQA solution, the camera continuously captures video streams from the user side and inputs the video streams to the visual big model. The microphone continuously captures audio from the user side and uploads the audio streams to the automatic speech recognition module. The automatic speech recognition module converts the audio streams into text, which is then input to the visual big model. The visual big model answers the user's questions based on the video streams and text. The answers are output in text form, and the text-to-speech module converts the text answers into speech, which is then transmitted to the player for playback.

[0072] As shown above, video VQA solutions require cameras to transmit images in real-time at a fixed frame rate and resolution, while large visual models need to process the video stream in real-time. This places excessive demands on the real-time computing power of the visual model. Currently, the resolution of smart terminals is generally above 1080P, even reaching 4K, with frame rates of 15-30fps. This requires the real-time computing power of the large visual model to reach 1080P@15fps, but the computing power of current mainstream large visual models is far from meeting this requirement, resulting in the slow development of video VQA. Currently, question-answering systems on smart terminals focus on image VQA, not video VQA. Furthermore, due to the temporal redundancy of video, there are very few changes in the image between consecutive video frames. For large visual models, they only need to perceive these changes to perceive changes in the physical world. For static scenes, there may not even be any image changes between consecutive images, and large visual models use most of their computing power for analyzing repetitive scenes, resulting in a serious waste of computing power. Furthermore, for users, real-time performance and accuracy are the most important metrics for question-answering systems. However, large-scale visual models are constantly under heavy computational pressure, making it difficult to provide rapid feedback and resulting in poor real-time performance. In other words, video VQA requires complex real-time audio and video transmission links. Large-scale visual models require powerful real-time computing power to support the processing of real-time video streams, as well as real-time transmission and computation, leading to high costs. Consequently, visual question answering suffers from high latency, poor accuracy, and limited flexibility.

[0073] This specification provides an embodiment of a visual task processing method for smart terminals, which enables real-time visual question answering in video question answering scenarios while ensuring efficient utilization of computing power, thereby solving the problems of high latency and poor accuracy in existing systems. Specifically, in the image acquisition stage, image feature extraction is used to analyze the current pose and scene complexity, dynamically configuring the frame rate and resolution of the smart terminal's camera to reduce the input data to the camera; in the image uploading stage, image feature extraction is used to perform similarity comparison, uploading images with significant similarity variations; in the image usage stage, a small SLLM model is used to filter reference images fed into the large visual model. By reducing the number and resolution of images that the large visual model needs to analyze in each stage of question answering analysis, real-time visual question answering can be achieved in video question answering scenarios while ensuring efficient utilization of computing power, solving the problems of high latency and poor accuracy in existing visual question answering systems.

[0074] To address the aforementioned technical problems, this specification provides a visual task processing method for smart terminals. This specification also relates to a visual task processing system, an information processing method based on a visual processing model, a computing device, a smart terminal with a camera, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0075] Considering the large number of model parameters in large models and the limited computing resources of smart terminals, the visual task processing method for smart terminals provided in this application can be applied to, for example... Figure 3 The application architecture shown is not limited to this. In, for example... Figure 3 In the application architecture shown, the large visual model is deployed on server 10. Server 10 can connect to one or more smart terminals 20 via a local area network (LAN), wide area network (WAN), internet connection, or other types of data network. These smart terminals 20 may include, but are not limited to, wristbands, AI glasses, AI Rabbit, AI wearable devices, AR devices, smartphones, tablets, laptops, and PDAs. The smart terminals 20 can interact with the user through a graphical user interface to access the large visual model, thereby implementing the methods provided in the embodiments of this specification.

[0076] In the embodiments of this specification, the system consisting of a smart terminal and a server can perform the following steps: the smart terminal performs the following steps: "acquiring task parameters of the target visual task, wherein the task parameters include task information of the target visual task and a corresponding reference time range; filtering a target image sequence from a reference image set according to the reference time range, wherein the reference image set stores at least one frame of image acquired in chronological order; transmitting the target image sequence and task information to the server," and the server performs the following steps: "obtaining the target processing result of the target visual task through a target visual processing model based on the target image sequence and task information; returning the target processing result to the smart terminal." It should be noted that, provided that the operating resources of the smart terminal can meet the deployment and operating conditions of the large visual model, the embodiments of this application can be performed on the smart terminal.

[0077] See Figure 4 , Figure 4 A flowchart of a visual task processing method for a smart terminal according to an embodiment of this specification is shown, which specifically includes the following steps.

[0078] Step 402: Obtain the task parameters of the target visual task, wherein the task parameters include the task information of the target visual task and the corresponding reference time range, the reference time range being the time range prior to the target visual task.

[0079] The embodiments in this specification apply to applications, websites, or mini-programs with visual task processing capabilities. These applications, websites, or mini-programs implement visual task processing functions. For example, a mini-program that deploys a target visual processing model can implement visual task processing functions such as visual question answering, visual content search, and task instruction processing. Another example is a third-party application that can implement corresponding visual task processing functions by calling the deployed target visual processing model through an Application Programming Interface (API).

[0080] The target visual task is the task to be processed. This target visual task instructs visual processing based on task parameters. For example, the target visual task could be a video-based question-and-answer task (e.g., what brand of car was just passed?), a video-based content search task (e.g., searching for other parks near the park visited yesterday), or a video-based task instruction processing task (e.g., navigating to the park visited yesterday). The task parameters can include task information and a corresponding reference time range. Task information refers to the detailed task content to be processed, such as the question or search keywords. The reference time range refers to the time range preceding the target visual task. This reference time range can be related to the time the target visual task was initiated, such as currently, just now, or 10 minutes ago, or it can be unrelated to the time the target visual task was initiated, such as yesterday, January 10th (currently January 12th), or last month. For example, if the target visual task information is "what brand of car was just passed?", the corresponding reference time range is a period before the target visual task was initiated; if the target visual task information is "what is the name of the park I arrived at at 10 am yesterday?", the corresponding reference time range is a period before and after 10 am yesterday.

[0081] In an optional implementation of this embodiment, the task parameters for the target visual task are obtained, including:

[0082] The task voice is collected using a voice acquisition device, and the start and end times of the task voice are obtained.

[0083] Obtain task information based on task voice prompts;

[0084] Based on the start time, end time, and task information, determine the reference time range corresponding to the target visual task.

[0085] In practice, smart devices are equipped with a voice acquisition device, which can be a microphone or a microphone array. The voice acquisition device can collect the user-initiated task voice and can distinguish between voice and non-voice through sound VAD technology to determine the start and end times of the user-initiated task voice.

[0086] Specifically, based on sound 3A technology, the acquired task speech can be preprocessed with noise reduction, volume normalization, and removal of silent segments to obtain the audio to be processed. Then, based on automatic speech recognition, the acquired audio to be processed is converted into corresponding text. Specifically, audio features, such as spectrograms and Mel-frequency cepstral coefficients (MFCCs), can be extracted from the audio to represent different attributes of the speech. Then, a pre-trained machine learning or deep learning model is used to analyze these audio features and attempt to match the most likely word or phrase sequences. Acoustic models and language models are typically used. The acoustic model is used to predict the phoneme probability distribution corresponding to a given audio segment, and the language model helps to determine which word combinations are most likely to occur in natural language, thereby improving the reasonableness of the transcription results. One or more possible texts are output. Based on the confidence score, the option with the higher probability is selected as the text obtained by recognition. This text is the task information corresponding to the task speech.

[0087] It should be noted that start and end times represent the beginning and end of the task's audio, while task information represents the specific detailed task content. Analyzing the start and end times, task information, and task intent allows for the inference of the task's intent and the determination of the corresponding reference time range.

[0088] For example, if a user starts saying "What brand was the car that just passed by?" at 10:00 and ends at 10:02, we can determine the start time of the task voice is "10:00", the end time is "10:02", and the task information is the text "What brand was the car that just passed by?". Based on the start time "10:00", the end time "10:02", and the task information "What brand was the car that just passed by?", we can determine the corresponding reference time range. Similarly, if a user starts saying "What was the name of the park I arrived at at 10:00 yesterday morning?" at 15:00 and ends at 15:05, we can determine the start time of the task voice is "15:00", the end time is "15:05", and the task information is the text "What was the name of the park I arrived at at 10:00 yesterday morning?". Based on the start time "15:00", the end time "15:05", and the task information "What was the name of the park I arrived at at 10:00 yesterday morning?", we can determine the corresponding reference time range.

[0089] In the embodiments of this specification, a voice acquisition device can collect the user's voice in real time to obtain the task voice, record the start and end times of the task voice, and then convert the task voice into text to obtain the corresponding task information. The start and end times and task information are then inferred to determine the corresponding reference time range. The obtained task information and reference time range are used as task parameters for the target visual task. The user can directly state the task information of the visual task to be performed. The smart terminal can collect the task voice, automatically identify the corresponding task information and the corresponding reference time range, facilitating user operation and enabling the smart terminal to automatically analyze the task information and reference time range to achieve the processing of the target visual task.

[0090] Of course, in actual implementation, in addition to initiating the target visual task via voice as described above, users can also directly input task text as task information through a smart terminal, collect the start and end times of input as start and end times, and thus determine the reference time range corresponding to the target visual task. This specification does not limit this.

[0091] In one optional implementation of this embodiment, a reference time range corresponding to the target visual task is determined based on the start time, end time, and task information, including:

[0092] Perform semantic analysis on the task information to obtain the time information included in the task information;

[0093] If the time information is related to the task time, then a reference time range is determined based on the start time and / or end time;

[0094] If the time information is not related to the task time, then the reference time range is determined based on the historical time indicated by the task information.

[0095] In practice, the task information includes at least one sentence. Each sentence in the task information is segmented into words, and each word is assigned a part-of-speech tag (noun, verb, adjective, etc.). All text is converted to a uniform format, such as all lowercase, and punctuation is removed, simplifying subsequent processing. Then, a pre-trained model or custom rules are used to identify named entities in the text. For extracting time information, time-related entities can be found, such as "just now," "yesterday," "March 5th," "2 PM," and "tomorrow." Specifically, common time expressions can be captured by writing specific pattern matching rules (regular expressions), or by using models trained on a large amount of labeled data to automatically identify and classify time expressions. These models may be based on Support Vector Machines (SVM), Conditional Random Fields (CRF), Recurrent Neural Networks (RNN), or Transformer architectures, etc. After identifying time-related entities, they can be converted into a standard time format (e.g., ISO 8601). Different processing methods can be considered for relative times (e.g., "yesterday," "next Wednesday") and standard times (e.g., "April 1, 2023"). Specifically, the actual date of the relative time can be inferred from the context of the sentence. For example, "next Friday" requires knowing the current date to determine which Friday it is. If cross-timezone scheduling is involved, time zone differences also need to be considered to ensure the accuracy of the time information. Afterward, the obtained time expressions are transformed into a unified form that is easy for computers to process.

[0096] It should be noted that "related to task time" refers to being related to the current time of the target visual task, while "unrelated to task time" means being unrelated to the current time of the target visual task but related to a certain historical time. Through the semantic analysis process described above, the time information included in the task information can be obtained. If this time information is related to the task time, a reference time range is determined based on the start and / or end times. For example, if the time information is related to the start time of the target visual task, the set time range before the start time of the target visual task can be used as the corresponding reference time range. If the time information is related to the end time of the target visual task, the set time range before the end time of the target visual task can be used as the corresponding reference time range. If the time information is unrelated to the task time, the historical time indicated by the task information can be determined to establish a reference time range. For example, if the time information indicates a historical time, the set time range before and / or after the historical time can be used as the reference time range.

[0097] Using the previous example, if the task information is the text "What brand was the car that just passed by?", and the recognized time information is "just now", which is related to the task time, then the period before the start time "10:00" can be used to determine the corresponding reference time range. If the task information is the text "What was the name of the park I arrived at at 10:00 AM yesterday?", and the recognized time information is "10:00 AM yesterday", then the period before and after "10:00 AM yesterday" can be used to determine the reference time range.

[0098] In the embodiments of this specification, the time information in the task information of the target visual task can be determined, and the time information can be distinguished between those related to the task time and those unrelated to the task time. Different methods are used to determine the corresponding reference time range, which can adapt to a variety of different types of task scenarios, making it more flexible and more accurate in determining the corresponding reference time range, thus ensuring the accuracy of visual task processing.

[0099] In one optional implementation of this embodiment, a reference time range corresponding to the target visual task is determined based on the start time, end time, and task information, including:

[0100] The start time, end time, and task information are input into the language processing model to obtain the reference time range output by the language processing model. The model size of the language processing model is smaller than a set size threshold. The language processing model is trained based on the start time, end time, and task information of the sample visual task, as well as the corresponding time range labels.

[0101] In practice, an initial semantic model can be pre-trained using supervised training based on the start and end times of the sample visual task, task information, and corresponding time range labels. This training yields a language processing model whose size is smaller than a set size threshold, which can be customized based on actual needs; that is, the language processing model is a small-scale language model (SLLM). The start and end times, and task information are input into the language processing model. The model analyzes the task information and, combined with the start and end times of the target visual task, outputs a reference time range of image frames required for the large-scale visual model to process the target visual task.

[0102] Using the previous example, inputting the start time "10:00", end time "10:02", and task information "What brand was the car that just passed by?" into the trained language processing model will output a reference time range of "9:55-10:00". Inputting the start time "15:00", end time "15:05", and task information "What was the name of the park I arrived at at 10:00 yesterday morning?" into the trained language processing model will output a reference time range of "October 11 (assuming the current time is October 12) 9:50-10:10".

[0103] In the embodiments of this specification, a small-sized language processing model can be directly used to analyze the start time, end time, and task information to obtain a reference time range. Since the language processing model is a small-sized model, the processing speed is fast, which greatly improves the efficiency of determining the reference time range. It can quickly obtain the reference time range of the image frames required for the large visual model to process the target visual task, which is convenient for quickly filtering out the image frames within the corresponding time range for the large visual model to process the task.

[0104] Step 404: Select the target image sequence from the reference image set according to the reference time range. The reference image set stores at least one frame of image acquired in chronological order before the target visual task. The target image sequence is acquired by a camera on the smart terminal that matches the target visual task.

[0105] Specifically, the reference image set is obtained based on the video stream continuously collected by the smart terminal. That is to say, the smart terminal can continuously collect video streams through the camera and write them into the reference image set. After receiving the target vision task, the corresponding target image sequence can be selected from the reference image set based on the determined reference time range for subsequent target vision task processing.

[0106] In one optional implementation of this embodiment, selecting the target image sequence from the reference image set based on a reference time range includes:

[0107] Determine the start and end timestamps based on the reference time range;

[0108] Images between the start and end timestamps are selected from the images in the reference image set as the target image sequence.

[0109] In practice, each image in the reference image set carries a corresponding timestamp, allowing the determination of the reference time range, including the start and end timestamps. Images between these timestamps are then selected as the target image sequence. Specifically, images in the reference image set typically contain EXIF ​​(Exchangeable Image File Format) or other metadata, which may include the capture time. By reading this information, the original timestamps of the images can be obtained. Using the start and end timestamps as boundary conditions, images falling between these boundaries are selected. The selected images are then sorted according to their timestamps, ensuring they are arranged in chronological order. Finally, the selected images are organized chronologically into a continuous sequence, forming the target image sequence.

[0110] Of course, in actual implementation, in addition to selecting images located between the start time stamp and the end time stamp as the target image sequence, the target image sequence can also be selected based on a range set before / after the start time stamp and a range set before / after the end time stamp. That is, the selected target image sequence can be the same as the reference time range or it can be different.

[0111] In the embodiments of this specification, the intelligent terminal can continuously collect video streams. When a target visual task needs to be processed, it can filter out the target image sequence between the start and end timestamps of the reference time range of the target visual task. Subsequently, the target visual task can be processed based on the filtered target image sequence. Only the image sequence related to the target visual task is processed, without the need to process each frame of the video stream in real time. This reduces the number of images that need to be processed, saves model computing power, and ensures the processing efficiency of the visual task.

[0112] In an optional implementation of this embodiment, the method further includes:

[0113] Image acquisition is performed based on the target acquisition parameters to obtain candidate images, where the target acquisition parameters are the current image acquisition parameters;

[0114] Determine the similarity between candidate images and images in the reference image set;

[0115] If the similarity is less than the similarity threshold, the reference image set is updated based on the candidate images.

[0116] It should be noted that the smart terminal can continuously capture video streams and write them into a reference image set as task reference data for the target visual task. In actual implementation, the smart terminal's camera can be controlled to capture images according to the target acquisition parameters to obtain candidate images. The target acquisition parameters are the current image acquisition parameters, which can be the acquisition frame rate, resolution, etc.

[0117] In practice, due to the temporal redundancy of video, there is little change between consecutive image frames. For static scenes, there may even be no change between consecutive image frames. Therefore, the similarity between a candidate image and an image in the reference image set can be determined. If the similarity is less than a similarity threshold, it indicates that the candidate image and the image in the reference image set have low similarity, and the candidate image is added to the reference image set to update the reference images. If the similarity is greater than or equal to the similarity threshold, it indicates that the candidate image and the image in the reference image set have high similarity, and the candidate image can be discarded to avoid redundancy. The similarity threshold is a pre-configured value used to determine the degree of similarity between two images.

[0118] In practical implementation, since the temporal redundancy of video is often time-related, the similarity between the currently obtained candidate image and the previous frame image in the reference image set can be calculated. The previous frame image in the reference image set refers to the image in the reference image set whose timestamp is closest to the candidate image, i.e., the last image written into the reference image set before the current time. This allows for the calculation of the similarity between consecutive image frames. Specifically, image features of the candidate image and the previous frame image in the reference image set can be extracted. An image similarity detection algorithm is then used to detect the image similarity features between the candidate image and the previous frame image in the reference image set, thus obtaining the similarity between the two images. Of course, in actual implementation, the similarity between the currently obtained candidate image and other images in the reference image set can also be calculated, such as the previous two frames, or each image frame within a certain time range. This specification does not limit this approach.

[0119] In the embodiments of this specification, the similarity between the currently obtained candidate image and the images in the reference image set can be calculated. Candidate images with similarity less than the similarity threshold are uploaded and written into the reference image set. The reference image set only stores consecutive images with low similarity. That is, during the image uploading stage, the similarity between consecutive images is compared through image feature extraction. Images with large similarity changes are uploaded to reduce the number of images stored in the reference image set, thereby reducing the number of images that need to be processed by the target vision processing model in the future and ensuring efficient use of computing power.

[0120] In one optional implementation of this embodiment, image acquisition is performed based on target acquisition parameters to obtain candidate images, including:

[0121] Image acquisition is performed based on the target acquisition parameters, and the quality parameters of the currently acquired image are determined.

[0122] If the quality parameters meet the quality constraints, the currently acquired image is used as a candidate image.

[0123] It should be noted that image acquisition can be performed based on target acquisition parameters to determine the quality parameters of the acquired image. These quality parameters reflect the quality of the acquired image and can include image sharpness, contrast, color accuracy, resolution, etc. Quality constraints refer to the constraints corresponding to the quality parameters, used to determine whether the quality parameters of the acquired image meet the quality requirements. If the quality constraints are met, it means that the acquired image is of good quality and can be used for subsequent target vision processing model tasks. If the quality constraints are not met, it means that the acquired image is of poor quality and cannot be used for subsequent target vision processing model tasks.

[0124] In practical implementation, image clarity is a fundamental condition for the target visual processing model to perform image analysis. Insufficient image clarity not only prevents the perception of specific image content but also introduces noise. Therefore, taking sharpness as a quality parameter, the quality constraint can be that the sharpness is higher than a set sharpness threshold. Specifically, after acquiring the image based on the target acquisition parameters, an image sharpness detection algorithm can be used to determine the image sharpness characteristics of the currently acquired image. Based on these characteristics, the sharpness of the currently acquired image is determined. If the sharpness is higher than the set sharpness threshold, the currently acquired image is used as a candidate image for further filtering and judgment. If the sharpness is lower than or equal to the set sharpness threshold, it indicates that the currently acquired image has poor sharpness and is directly discarded without further filtering and judgment.

[0125] In the embodiments of this specification, after image acquisition is performed according to the target acquisition parameters, the quality parameters of the currently acquired image can be further determined. If the quality parameters meet the quality constraints, the currently acquired image is then used as a candidate image for subsequent similarity analysis, model reasoning, and other processes. This avoids misjudgment or waste of computing power caused by analyzing images with poor quality, and improves the accuracy and efficiency of visual task processing.

[0126] In an optional implementation of this embodiment, after acquiring candidate images based on target acquisition parameters, the method further includes:

[0127] Determine the image complexity of candidate images and obtain historical image acquisition information;

[0128] Based on image complexity and historical image acquisition information, the target acquisition parameters of the smart terminal are determined.

[0129] In practical implementation, after acquiring candidate images based on target acquisition parameters, an image complexity detection algorithm can be used to determine the image complexity features of the candidate images. Based on these features, the image complexity of the candidate image is determined. Image complexity can indicate the complexity of the currently captured scene, such as simple, ordinary, or complex. Then, based on the determined image complexity and historical image acquisition information, the target acquisition parameters of the smart terminal can be dynamically adjusted. Historical image acquisition information can refer to historical camera parameters, historical poses, historical complexity, etc.

[0130] It should be noted that, taking target acquisition parameters including resolution as an example, the more complex the image, the higher the required resolution. Therefore, dynamic update rules for target acquisition parameters can be pre-configured, such as increasing the resolution when the complexity increases, or increasing the resolution to the set value when the complexity reaches a set level.

[0131] As an example, it can be determined whether the complexity of the current candidate image has increased compared to the historical complexity. If so, the historical resolution is determined based on the historical camera parameters. If the historical resolution is low, the set resolution can be increased on the historical resolution to obtain the updated resolution as the target acquisition parameter of the smart terminal. Subsequently, images are continuously acquired with the dynamically updated resolution.

[0132] In the embodiments of this specification, the target acquisition parameters of the smart terminal can be dynamically updated based on the image complexity of the candidate image and combined with historical image acquisition information. For example, if the image complexity is high, the resolution is increased, and if the image complexity is low, the resolution is decreased. This allows the acquired images to adapt to the complexity of the actual scene, ensuring the subsequent image processing and analysis effects, while avoiding the waste of resources caused by always acquiring images at a high resolution. To a certain extent, this reduces the resolution of images that need to be processed by the large visual model, ensuring the real-time computing power of the model.

[0133] In an optional implementation of this embodiment, before determining the target acquisition parameters of the smart terminal based on image complexity and historical image acquisition information, the following steps are also included:

[0134] Collect the current attitude parameters of the smart terminal, including the speed parameters of the smart terminal;

[0135] Accordingly, based on image complexity and historical image acquisition information, the target acquisition parameters of the smart terminal are determined, including:

[0136] The target acquisition parameters of the smart terminal are determined based on image complexity, current posture parameters, and historical image acquisition information.

[0137] In practical implementation, if the smart terminal is a mobile device, its current posture parameters can also be collected. These parameters include the terminal's velocity, which reflects the device's motion type. By combining image complexity, current posture parameters, and historical image acquisition information, the target acquisition parameters for the smart terminal can be determined. Specifically, the current posture parameters can be collected using the posture sensor configured on the smart terminal. After processing such as jitter reduction and normalization, motion type analysis can be performed. For example, by collecting the smart terminal's acceleration information using an accelerometer, and processing it with jitter reduction and normalization, the motion type of the smart terminal can be determined, such as stationary, slow, normal, or fast.

[0138] It should be noted that, taking the target acquisition parameters, including the frame rate, as an example, the faster the smart terminal moves, the higher the required frame rate. Therefore, dynamic update rules for the target acquisition parameters can be pre-configured, such as increasing the acquisition frame rate when the motion type becomes faster, or increasing the frame rate to the set value when the motion type reaches a set level.

[0139] As an example, it can be determined whether the complexity of the current candidate image has increased compared to the historical complexity. If so, the historical resolution is determined based on the historical camera parameters. If the historical resolution is low, a set resolution can be added to the historical resolution to obtain the updated resolution. Furthermore, the motion type indicated by the current posture parameters of the smart terminal can be determined, and the frame rate corresponding to the motion type can be obtained as the updated frame rate. The dynamically determined updated resolution and updated frame rate are used as the target acquisition parameters of the smart terminal, and images are continuously acquired subsequently with the dynamically updated resolution and frame rate.

[0140] For example, if the current smart terminal's motion type is stationary and the image complexity is simple, the frame rate and resolution can be gradually reduced, or even image acquisition can be stopped; if the current smart terminal's motion type changes from stationary to slow, the frame rate can be gradually increased, and if the image complexity becomes simpler, the resolution can be gradually reduced.

[0141] In the embodiments of this specification, the target acquisition parameters of the smart terminal can be dynamically updated based on the image complexity of the candidate image, the current posture parameters of the smart terminal, and historical image acquisition information. For example, if the image complexity is high, the resolution is increased; if the movement is fast, the frame rate is increased; if the image complexity is low, the resolution is decreased; if the movement is slow, the frame rate is decreased. This allows the acquired images to adapt to the complexity of the actual scene and the movement type of the smart terminal, ensuring the effectiveness of subsequent image processing and analysis. It also avoids the waste of resources caused by always acquiring images at a high resolution and high frame rate. By integrating and analyzing image complexity, current posture parameters, and historical image acquisition information, the acquisition resolution and frame rate of the smart terminal can be dynamically updated. This reduces the resolution and number of images that need to be processed by the large visual model to a certain extent, ensuring the real-time computing power of the model.

[0142] Step 406: Based on the target image sequence and task information, obtain the target processing result of the target visual task through the target visual processing model.

[0143] It should be noted that the task information of the target vision task and the target image sequence corresponding to the reference time range selected from the reference image set can be transmitted to the target vision processing model. The target vision processing model can then perform inference analysis on the input target image sequence to obtain the result corresponding to the task information, which serves as the target processing result of the target vision task.

[0144] The target visual processing model is a pre-trained large-scale visual model capable of analyzing visual content (such as images and videos) and providing corresponding inference results. In one implementation, the target visual processing model is deployed on a server, and the smart terminal can call the deployed target visual processing model through an interface to analyze the target image sequence and task information, outputting the corresponding target processing results. In another implementation, if the smart terminal's operating resources can meet the deployment and operating conditions of the target visual processing model, the target visual processing model can also be directly deployed on the smart terminal to achieve visual task processing.

[0145] In practice, taking visual question answering as an example of a target visual task, the task information is the question to be answered. The target image sequence and the question to be answered can be input into the target visual processing model. The target visual processing model analyzes the input target image sequence, determines the answer to the question to be answered, and outputs it.

[0146] Specifically, the target visual processing model outputs the target visual task processing results in text form, which can be directly output to the display device of a smart terminal for feedback to the user. Alternatively, text-to-speech (TTS) technology can be used to convert the text-based target processing results into speech form, which can then be played through the smart terminal's player, allowing the user to directly hear the target visual task processing results for easier operation.

[0147] This specification provides a visual task processing method for smart terminals. When processing visual tasks, a reference time range can be quickly determined using a language processing model (a large language model with relatively small parameters). The target image sequence corresponding to the reference time range is selected from at least one frame of images acquired chronologically prior to the target visual task, serving as reference information for processing the visual task. This ensures real-time image processing while reducing the number of images the target visual processing model needs to process, thus guaranteeing both processing efficiency and real-time performance, and placing lower demands on the real-time computing power of the target visual processing model. Furthermore, the posture acquisition parameters of the smart terminal can be dynamically updated based on image complexity, the smart terminal's posture parameters, and historical image acquisition information. This reduces the acquisition resolution and frame rate while maintaining image processing quality, further saving model computing power. Moreover, task processing can be implemented based on the target image sequence corresponding to the reference time range, leveraging rich image information to improve the accuracy of visual task processing.

[0148] Figure 5 This is a schematic diagram of the processing procedure of a vision task processing system provided in one embodiment of this specification, as shown below. Figure 5 As shown, the visual task processing system is deployed on a smart terminal and includes a preprocessing module, a dynamic configuration module, an image uploading module, and an intelligent filtering module.

[0149] Images are continuously captured by the camera of the smart terminal. The preprocessing module extracts image sharpness features from the captured images to determine whether the sharpness meets the set sharpness threshold. If not, the image is discarded; if so, the image is clear enough and can be used as a candidate image for further analysis and processing in the subsequent dynamic configuration module and image upload module.

[0150] The image upload module can extract image similarity features from candidate images, calculate the image similarity between the candidate image and the previous frame image in the reference image set. If the similarity is less than the similarity threshold, the candidate image is uploaded to the Object Storage Service (OSS) and written into the reference image set; if the similarity is greater than or equal to the similarity threshold, the candidate image is discarded and not uploaded.

[0151] Additionally, current posture parameters can be collected via the smart terminal's posture sensor. A pre-processing module can then perform data processing such as jitter reduction and normalization on these parameters before transmitting them to the dynamic configuration module. The dynamic configuration module performs posture analysis to determine the smart terminal's motion type (stationary, slow, normal, fast). Furthermore, the dynamic configuration module can acquire candidate images transmitted from the pre-processing module, extract image complexity features, and determine the image complexity (simple, normal, complex). Subsequently, by combining motion type (stationary, slow, normal, fast), image complexity (simple, normal, complex), and historical image acquisition information (historical camera parameters, historical posture, historical complexity), a fusion analysis can be performed to dynamically configure the smart terminal's target acquisition parameters (resolution, frame rate). Based on these target acquisition parameters, the camera is controlled to continuously acquire images.

[0152] Audio is continuously collected via the microphone of the smart terminal to obtain task speech. The Voice Activity Detection (VAD) module of the intelligent filtering module distinguishes between speech and non-speech, determining the start and end times of the user-initiated task speech. Then, the Automatic Speech Recognition (ASR) module of the intelligent filtering module converts the task speech into text-based task information. The start and end times of the task speech and the text-based task information are input into a Language Processing Model (SLLM) to obtain a corresponding reference time range. This SLLM is a large language model with relatively small parameters, suitable for running on smart terminals with low computing power, and offers good real-time performance. Based on this reference time range, the corresponding target image sequence is selected from the reference image set stored in the object storage service. This target image sequence and the task information are transmitted to a Visual Large Model (VL model) for target visual task processing, obtaining the corresponding text-based target processing result. This target processing result is then converted into audio-based target processing result using Text-to-Speech (TTS) technology, and the audio-based target processing result is fed back to the smart terminal's player for playback.

[0153] It should be noted that the aforementioned external components such as attitude sensors, cameras, microphones, and large-scale vision models can be standard modules or configured freely by the user.

[0154] In the embodiments of this specification, the visual task processing system is deployed on a smart terminal to implement a visual task processing method. During visual task processing, a reference time range can be quickly determined using a language processing model (a large language model with relatively small parameters). A target image sequence corresponding to the reference time range is selected from at least one frame of images acquired in chronological order as reference information for processing the visual task. This ensures the real-time performance of the images and reduces the number of images that the target visual processing model needs to process, guaranteeing both the processing efficiency and real-time performance of the visual task, while also lowering the real-time computing power requirements of the target visual processing model. Furthermore, the posture acquisition parameters of the smart terminal can be dynamically updated based on image complexity, the smart terminal's posture parameters, and historical image acquisition information. This reduces the acquisition resolution and frame rate while maintaining image processing effectiveness, further saving model computing power. Moreover, task processing can be implemented based on the target image sequence corresponding to the reference time range, referencing rich image information, improving the accuracy of visual task processing, and ensuring the real-time performance of the visual question-answering system.

[0155] The following is in conjunction with the appendix Figure 6 Taking the application of the visual task processing method provided in this manual to a visual question-answering task as an example, the visual task processing method will be further explained. Figure 6 This specification illustrates a flowchart of a visual task processing method for a smart terminal according to an embodiment of the present specification, which specifically includes the following steps.

[0156] Step 602: Obtain the current posture parameters of the smart terminal, determine the motion type of the smart terminal based on the current posture parameters; determine the image complexity of the candidate images; obtain historical image acquisition information; determine the updated target acquisition parameters of the smart terminal based on the image complexity, current posture parameters, and historical image acquisition information, and continue image acquisition based on the updated target acquisition parameters.

[0157] Step 604: Perform image acquisition based on the current target acquisition parameters to obtain the currently acquired image; determine the image clarity of the currently acquired image. If the image clarity meets the clarity threshold, then the currently acquired image is used as a candidate image.

[0158] Step 606: Determine the similarity between the candidate image and the images in the reference image set; if the similarity is less than the similarity threshold, then write the candidate image into the reference image set.

[0159] Step 608: Receive the user's question audio through the microphone and determine the initiation and end times of the question; convert the question audio into question text; based on the initiation time, end time, and question text, use a language processing model to determine a reference time range; based on the reference time range, select the target image sequence corresponding to the range from the reference image set.

[0160] Step 610: Input the target image sequence and question text into the question-answering model, and obtain the answer text corresponding to the question through the question-answering model.

[0161] Step 612: Convert the answer text into answer audio and play the answer audio through the player on the smart terminal.

[0162] This specification provides a visual task processing method for smart terminals. When performing visual question answering, a reference time range can be quickly determined using a language processing model (a large language model with relatively small parameters). Target image sequences corresponding to the reference time range are selected from at least one frame of images acquired in chronological order as reference information for visual question answering. This ensures real-time image processing while reducing the number of images that the large question answering model needs to process, thus guaranteeing both processing efficiency and real-time performance, and placing lower demands on the real-time computing power of the large question answering model. Furthermore, the posture acquisition parameters of the smart terminal can be dynamically updated based on image complexity, the smart terminal's posture parameters, and historical image acquisition information. This reduces the acquisition resolution and frame rate while maintaining image processing quality, further saving model computing power. Moreover, visual question answering can be achieved based on the target image sequences corresponding to the reference time range, referencing rich image information and improving the accuracy of visual question answering.

[0163] See Figure 7 , Figure 7 A flowchart of an information processing method based on a vision processing model according to an embodiment of this specification is shown, which is applied to a task platform and specifically includes the following steps 702-704.

[0164] Step 702: Receive a model request sent by the smart terminal, wherein the model request includes at least one of the following: scene identifier of the target scene, scene input data of the target scene, and model specification parameters.

[0165] Step 704: Based on the model request, determine the corresponding target visual processing model from at least one visual processing model, wherein the target visual processing model is used to execute the above-described visual task processing method applied to the smart terminal.

[0166] It should be noted that, based on the model request, the corresponding target visual processing model is determined from at least one visual processing model. One possible approach is to search for the corresponding target visual processing model from at least one visual processing model included in the model library based on the model request; another possible approach is to train the target visual processing model based on the model request; yet another possible approach is to construct the target visual processing model based on the model request, which is not limited here.

[0167] For example, based on the scene identifier of the target scene, at least one pre-trained visual processing model can be found in the model library. Then, based on the model specification parameters, a visual processing model of the corresponding size can be selected from the at least one visual processing model. Finally, based on the scene input data of the target scene, the visual processing model of the corresponding size can be trained to obtain a target visual processing model suitable for user needs.

[0168] In one optional implementation of this embodiment, the model request includes a scene identifier of the target scene; based on the model request, determining the corresponding target visual processing model from at least one visual processing model includes:

[0169] Based on the scene identifier of the target scene, a target visual processing model suitable for the target scene is searched from the model library. The model library stores at least one visual processing model suitable for different visual processing scenes.

[0170] It should be noted that the model library is a database for storing and managing various pre-trained deep learning models. Multiple vision processing models adapted to different vision processing scenarios cover different application scenarios and needs. The model library allows users to select the appropriate model according to their needs, or directly use the model for vision task processing through API calls.

[0171] Multiple visual processing models adapted to different visual processing scenarios are stored in the model library, each specifically designed for different visual processing scenarios. Each model is optimized for a specific application environment, and any given visual processing model is trained using the aforementioned visual processing model training method. For example, based on the scene identifier "video question answering generation," a target visual processing model suitable for the video question answering scenario can be found in the model library.

[0172] In the embodiments described in this specification, based on scene requirements, the target visual processing model adapted to the scene is accurately found through scene identification, making the target processing results more accurate and scene-appropriate, thereby improving user experience and visual task processing quality.

[0173] As an example, the task platform can provide target visual processing models for various scenarios. For instance, in a video question-and-answer scenario, it can provide the corresponding target visual processing model based on the model request sent by the smart terminal, thereby enabling the corresponding visual task processing.

[0174] In one optional implementation of this embodiment, the model request includes scene input data of the target scene; based on the model request, determining the corresponding target visual processing model from at least one visual processing model includes:

[0175] From at least one visual processing model, determine an initial visual processing model that is suitable for the target scene;

[0176] Based on the scene input data of the target scene, the initial visual processing model is trained to obtain the target visual processing model.

[0177] In actual implementation, the model request may include scene input data of the target scene, and the target visual processing model is a visual processing model suitable for the target scene.

[0178] For example, a general visual processing model is a basic visual processing model that is trained to adapt to different visual processing scenarios, but is not optimized for any specific scenario. For instance, based on scene input data from a video question-answering scenario, the general visual processing model can be trained to obtain a target visual processing model adapted to the video question-answering scenario.

[0179] In the embodiments described in this specification, based on scenario requirements, a general visual processing model is further trained using scenario input data to obtain a target visual processing model adapted to the scenario, making the target processing results more accurate and scenario-appropriate, thereby improving user experience and the processing quality of visual tasks.

[0180] In one optional implementation of this embodiment, the model request includes model specification parameters; based on the model request, determining the corresponding target visual processing model from at least one visual processing model includes:

[0181] Based on the model specification parameters, the corresponding target visual processing model is searched from the model library, which stores multiple visual processing models with different model specification parameters.

[0182] Among them, the model specification parameter can be the model size, such as based on the model size: 32GB, to find the target visual processing model of the corresponding size from the model library.

[0183] In the embodiments described in this specification, based on the model specification requirements, the corresponding target visual processing model is accurately found through the model specification parameters, which ensures the efficient and stable operation of the target visual processing model and improves the user experience.

[0184] In an optional implementation of this embodiment, after determining the corresponding target visual processing model from at least one visual processing model based on the model request, the method further includes:

[0185] Deploy the target visual processing model and build a visual processing interface based on the target visual processing model so that the smart terminal can schedule the target visual processing model to execute visual tasks.

[0186] It should be noted that the visual processing interface is an interactive programming interface for smart terminals to schedule target visual processing models, typically provided in the form of an API. Through the visual processing interface, users can input task data for the target visual task, such as task information and the corresponding reference time range, and effectively control the model's output, such as outputting the corresponding answer.

[0187] In practical implementation, one possible approach to deploying the target visual processing model is to deploy it on a distributed system of the task platform. For example, the target visual processing model can be deployed on a distributed system of the task platform, and a visual processing interface can be built based on the target visual processing model and provided to the smart terminal, so that the smart terminal can schedule the target visual processing model to execute the target visual tasks in the visual processing scenario.

[0188] In the embodiments described in this specification, efficient terminal access is achieved, visual task processing is optimized, and the processing quality and response speed of visual tasks are improved.

[0189] The information processing method based on the visual processing model provided in the embodiments of this specification can be adapted to user needs to obtain the target visual processing model, realize personalized model service, provide users with an efficient, flexible and easy-to-use model service method, and improve user experience.

[0190] Corresponding to the above method embodiments, this specification also provides task platform embodiments. Figure 8 A schematic diagram of the structure of a task platform provided in one embodiment of this specification is shown. Figure 8 As shown, the task platform 800 includes: a request interface 802 and a response unit 804;

[0191] Request interface 802 is used to receive model requests sent by smart terminals, wherein the model request includes at least one of the following: scene identifier of the target scene, scene input data of the target scene, and model specification parameters.

[0192] The response unit 804 is used to determine a corresponding target visual processing model from at least one visual processing model based on a model request, wherein the target visual processing model is used to execute the above-described visual task processing method applied to a smart terminal.

[0193] Optionally, the task platform also includes a visual processing interface, which is constructed based on the target visual processing model;

[0194] The visual processing interface is used by smart terminals to schedule and execute the aforementioned visual task processing methods.

[0195] Optionally, the model request includes a scene identifier for the target scene;

[0196] Response unit 804 is further used for:

[0197] Based on the scene identifier of the target scene, a target visual processing model suitable for the target scene is searched from the model library. The model library stores at least one visual processing model suitable for different visual processing scenes.

[0198] Optionally, the model request includes scene input data for the target scene;

[0199] Response unit 804 is further used for:

[0200] From at least one visual processing model, determine an initial visual processing model that is suitable for the target scene;

[0201] Based on the scene input data of the target scene, the initial visual processing model is trained to obtain the target visual processing model.

[0202] Optionally, the model request may include model specification parameters;

[0203] Response unit 804 is further used for:

[0204] Based on the model specification parameters, the corresponding target visual processing model is searched from the model library, which stores multiple visual processing models with different model specification parameters.

[0205] Optionally, the task platform also includes a deployment module, configured as follows:

[0206] Deploy the target visual processing model and build a visual processing interface based on the target visual processing model so that the smart terminal can schedule the target visual processing model to execute visual tasks.

[0207] In the embodiments described in this specification, the task platform adapts to user needs to obtain target visual processing models, realizes personalized model services, provides users with an efficient, flexible and easy-to-use model service platform, and improves user experience.

[0208] The above is an illustrative scheme of a task platform according to this embodiment. It should be noted that the technical solution of this task platform and the technical solution of the information processing method based on the vision processing model described above belong to the same concept. For details not described in detail in the technical solution of the task platform, please refer to the description of the technical solution of the information processing method based on the vision processing model described above.

[0209] Corresponding to the above method embodiments, this specification also provides embodiments of a vision task processing system. Figure 9 A schematic diagram of the structure of a vision task processing system according to one embodiment of this specification is shown. Figure 9 As shown, the system is deployed on a smart terminal and includes:

[0210] The task receiving module 902 is configured to acquire task parameters of the target visual task, wherein the task parameters include task information of the target visual task and the corresponding reference time range, the reference time range being the time range prior to the target visual task;

[0211] The image filtering module 904 is configured to filter a target image sequence from a reference image set according to a reference time range, wherein the reference image set stores at least one frame of image acquired in chronological order prior to the target visual task, and the target image sequence is acquired by a camera on a smart terminal that matches the target visual task.

[0212] The visual reasoning module 906 is configured to obtain the target processing result of the target visual task through the target visual processing model based on the target image sequence and task information.

[0213] Optionally, the system also includes an image update module, configured to:

[0214] Image acquisition is performed based on the target acquisition parameters to obtain candidate images, where the target acquisition parameters are the current image acquisition parameters;

[0215] Determine the similarity between candidate images and images in the reference image set;

[0216] If the similarity is less than the similarity threshold, the reference image set is updated based on the candidate images.

[0217] Optionally, the image update module is further configured as follows:

[0218] Image acquisition is performed based on the target acquisition parameters, and the quality parameters of the currently acquired image are determined.

[0219] If the quality parameters meet the quality constraints, the currently acquired image is used as a candidate image.

[0220] Optionally, the system also includes a dynamic configuration module, configured as follows:

[0221] Determine the image complexity of candidate images and obtain historical image acquisition information;

[0222] Based on image complexity and historical image acquisition information, the target acquisition parameters of the smart terminal are determined.

[0223] Optionally, the system also includes a front-end acquisition module, configured as follows:

[0224] Collect the current attitude parameters of the smart terminal, including the speed parameters of the smart terminal;

[0225] Accordingly, the dynamic configuration module is further configured as follows:

[0226] The target acquisition parameters of the smart terminal are determined based on image complexity, current posture parameters, and historical image acquisition information.

[0227] Optionally, the task receiving module 902 is further configured to:

[0228] The task voice is collected using a voice acquisition device, and the start and end times of the task voice are obtained.

[0229] Obtain task information based on task voice prompts;

[0230] Based on the start time, end time, and task information, determine the reference time range corresponding to the target visual task.

[0231] Optionally, the task receiving module 902 is further configured to:

[0232] Perform semantic analysis on the task information to obtain the time information included in the task information;

[0233] If the time information is related to the task time, then a reference time range is determined based on the start time and / or end time;

[0234] If the time information is not related to the task time, then the reference time range is determined based on the historical time indicated by the task information.

[0235] Optionally, the task receiving module 902 is further configured to:

[0236] The start time, end time, and task information are input into the language processing model to obtain the reference time range output by the language processing model. The model size of the language processing model is smaller than a set size threshold. The language processing model is trained based on the start time, end time, and task information of the sample visual task, as well as the corresponding time range labels.

[0237] Optionally, the image filtering module 904 is further configured to:

[0238] Determine the start and end timestamps based on the reference time range;

[0239] Images between the start and end timestamps are selected from the images in the reference image set as the target image sequence.

[0240] This specification provides a visual task processing system deployed on a smart terminal, including a task receiving module, an image filtering module, and a visual inference module. During visual task processing, a reference time range can be quickly determined using a language processing model (a large language model with relatively small parameters). Target image sequences corresponding to the reference time range are filtered from at least one frame of images acquired in chronological order, serving as reference information for processing the visual task. This ensures both real-time image processing and reduces the number of images the target visual processing model needs to process, guaranteeing both processing efficiency and real-time performance while minimizing the real-time computing power requirements of the target visual processing model. Furthermore, a dynamic configuration module is included, which can dynamically update the smart terminal's posture acquisition parameters based on image complexity, the smart terminal's posture parameters, and historical image acquisition information. This reduces the acquisition resolution and frame rate while maintaining image processing quality, further saving model computing power. Moreover, task processing can be implemented based on the target image sequences corresponding to the reference time range, leveraging rich image information to improve the accuracy of visual task processing.

[0241] The above is an illustrative scheme of a visual task processing system according to this embodiment. It should be noted that the technical solution of this visual task processing system and the technical solution of the visual task processing method applied to smart terminals described above belong to the same concept. For details not described in detail in the technical solution of the visual task processing system, please refer to the description of the technical solution of the visual task processing method applied to smart terminals described above.

[0242] Figure 10 A structural block diagram of a computing device provided in one embodiment of this specification is shown.

[0243] The computing device 1000 includes:

[0244] Memory 1010 and processor 1020;

[0245] The memory 1010 is used to store computer programs / instructions, and the processor 1020 is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor 1020, they implement the steps of the above-described visual task processing method applied to a smart terminal.

[0246] In one or more embodiments of this specification, the computing device 1000 can be understood as an integrated smart terminal, including but not limited to a server, desktop computer, PC (Personal Computer), all-in-one model machine, mobile phone, tablet computer or other portable smart terminal, etc., and the computing device may have the model in the above embodiments of this application pre-installed.

[0247] Specifically, the computing device 1000 can pre-install various types of models, including but not limited to models in natural language processing, visual processing, speech processing, code processing, and multimodal task processing, thus providing diverse model selection. In different product forms, the computing device 1000 can support one or more model usage methods, including but not limited to model training, model invocation, model fine-tuning, model deployment, model inference, and application. In some product forms, the computing device 1000 also supports model management, including but not limited to multi-type model management (supporting the management of discriminative, generative, and other types of models), model version control (supporting the control of different model versions), and model evaluation (evaluating model performance and effectiveness based on model evaluation tools). In other product forms, the computing device 1000 can also create applications based on models, providing API (Application Programming Interface) calling capabilities. Users can call models into created applications through the API interface, and application management tools are also provided to manage and monitor the applications.

[0248] Furthermore, the computing device 1000 may also include data management (supporting the creation and management of model tuning datasets), a training center (providing abundant training resources to help users learn artificial intelligence technologies), and basic control capabilities (providing enterprise-level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, it provides a comprehensive and integrated device for artificial intelligence development, training, deployment, and application.

[0249] Figure 11 This specification illustrates a structural block diagram of a smart terminal with a camera according to one embodiment. Figure 11 As shown, the smart terminal 1100 is deployed on the camera 1102, and also includes a voice acquisition device 1104, a visual task processing system 1106, and a player 1108.

[0250] The camera, 1102, is configured to capture at least one frame of image and write it to a reference image set;

[0251] The voice acquisition device 1104 is configured to acquire the voice of the task and generate the target visual task.

[0252] The visual task processing system 1106 is configured to acquire task parameters of a target visual task, wherein the task parameters include task information of the target visual task and a corresponding reference time range, the reference time range being the time range prior to the target visual task; select a target image sequence from a reference image set according to the reference time range; and obtain the target processing result of the target visual task through a target visual processing model based on the target image sequence and task information.

[0253] Player 1108 is configured to play the target processing results of the target visual task.

[0254] In one embodiment of this specification, the above-described components of the smart terminal 1100 and Figure 11 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 11 The block diagram of the smart terminal shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0255] The smart terminal 1100 can be any type of terminal device with a camera, including mobile computers or mobile smart terminals (e.g., tablets, personal digital assistants, laptops, notebooks, PDAs, etc.), mobile phones (e.g., smartphones), wearable smart terminals (e.g., AI glasses (smart glasses), smartwatches, AI wearable devices, AR devices, etc.) or other types of mobile smart terminals (e.g., AI Rabbit (AI robot)).

[0256] By applying the embodiments in this specification, images can be captured by the camera on the smart terminal, and in conjunction with voice input tasks, the smart terminal can provide the processing capability of target visual tasks through the target visual processing model.

[0257] The above is an illustrative scheme of a smart terminal according to this embodiment. It should be noted that the technical solution of this smart terminal and the technical solution of the visual task processing method applied to the smart terminal described above belong to the same concept. For details not described in detail in the technical solution of the smart terminal, please refer to the description of the technical solution of the visual task processing method applied to the smart terminal described above.

[0258] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described visual task processing method applied to a smart terminal.

[0259] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the visual task processing method applied to a smart terminal described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the visual task processing method applied to a smart terminal described above.

[0260] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described visual task processing method applied to a smart terminal.

[0261] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-described visual task processing method applied to a smart terminal belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-described visual task processing method applied to a smart terminal.

[0262] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0263] Computer programs / instructions include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. Computer-readable media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in computer-readable media can be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0264] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.

[0265] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0266] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A visual task processing method applied to a smart terminal, comprising: Obtain task parameters for the target visual task, wherein the task parameters include task information of the target visual task and a corresponding reference time range, and the reference time range is the time range prior to the target visual task; The target image sequence is selected from the reference image set according to the reference time range, wherein the reference image set stores at least one frame of image acquired in chronological order before the target visual task, and the target image sequence is acquired by a camera on the smart terminal that matches the target visual task; Based on the target image sequence and the task information, the target processing result of the target visual task is obtained through the target visual processing model.

2. The visual task processing method applied to a smart terminal according to claim 1, the method further includes: Image acquisition is performed based on target acquisition parameters to obtain candidate images, wherein the target acquisition parameters are the current image acquisition parameters; Determine the similarity between the candidate image and the images in the reference image set; If the similarity is less than the similarity threshold, the reference image set is updated based on the candidate image.

3. The visual task processing method applied to a smart terminal according to claim 2, wherein the step of acquiring images based on target acquisition parameters to obtain candidate images includes: Image acquisition is performed based on the target acquisition parameters, and the quality parameters of the currently acquired image are determined. If the quality parameters meet the quality constraints, the currently acquired image is used as the candidate image.

4. The visual task processing method applied to a smart terminal according to claim 2, after obtaining candidate images by acquiring images according to target acquisition parameters, further includes: Determine the image complexity of the candidate images and obtain historical image acquisition information; The target acquisition parameters of the smart terminal are determined based on the image complexity and the historical image acquisition information.

5. The visual task processing method applied to a smart terminal according to claim 4, further comprising, before determining the target acquisition parameters of the smart terminal based on the image complexity and the historical image acquisition information: The current attitude parameters of the smart terminal are collected, wherein the current attitude parameters include the speed parameters of the smart terminal; Accordingly, determining the target acquisition parameters of the smart terminal based on the image complexity and the historical image acquisition information includes: The target acquisition parameters of the smart terminal are determined based on the image complexity, the current posture parameters, and the historical image acquisition information.

6. The visual task processing method applied to a smart terminal according to claim 1, wherein obtaining the task parameters of the target visual task includes: The task voice is collected using a voice acquisition device, and the start and end times of the task voice are obtained. The task information is obtained based on the task voice; Based on the start time, the end time, and the task information, a reference time range corresponding to the target visual task is determined.

7. The visual task processing method for a smart terminal according to claim 6, wherein determining the reference time range corresponding to the target visual task based on the start time, the end time, and the task information includes: Semantic analysis is performed on the task information to obtain the time information included in the task information; If the time information is related to the task time, then the reference time range is determined based on the start time and / or the end time; If the time information is not related to the task time, the reference time range is determined based on the historical time indicated by the task information.

8. The visual task processing method for a smart terminal according to claim 6, wherein determining the reference time range corresponding to the target visual task based on the start time, the end time, and the task information includes: The start time, the end time, and the task information are input into the language processing model to obtain the reference time range output by the language processing model. The model size of the language processing model is smaller than a set size threshold. The language processing model is trained based on the start time, end time, and task information of the sample visual task, as well as the corresponding time range labels.

9. The visual task processing method for a smart terminal according to claim 1, wherein the step of selecting the target image sequence from the reference image set according to the reference time range includes: The start and end timestamps are determined based on the reference time range; Images between the start timestamp and the end timestamp are selected from the images in the reference image set as the target image sequence.

10. A visual task processing system, deployed on a smart terminal, comprising: The task receiving module is configured to acquire task parameters of a target visual task, wherein the task parameters include task information of the target visual task and a corresponding reference time range, and the reference time range is the time range prior to the target visual task. An image filtering module is configured to filter a target image sequence from a reference image set according to the reference time range, wherein the reference image set stores at least one frame of image acquired in chronological order prior to the target visual task, and the target image sequence is acquired by a camera on the smart terminal that matches the target visual task; The visual reasoning module is configured to obtain the target processing result of the target visual task through a target visual processing model based on the target image sequence and the task information.

11. An information processing method based on a visual processing model, applied to a task platform, comprising: Receive a model request sent by a smart terminal, wherein the model request includes at least one of the following: scene identifier of the target scene, scene input data of the target scene, and model specification parameters; Based on the model request, a corresponding target visual processing model is determined from at least one visual processing model, wherein the target visual processing model is used to execute the visual task processing method for smart terminals as described in any one of claims 1-9.

12. The method according to claim 11, wherein the model request includes a scene identifier of the target scene; and determining the corresponding target visual processing model from at least one visual processing model based on the model request includes: Based on the scene identifier of the target scene, a target visual processing model adapted to the target scene is searched from the model library, wherein the model library stores at least one visual processing model adapted to different visual processing scenes. The model request includes scene input data of the target scene; determining the corresponding target visual processing model from at least one visual processing model based on the model request includes: From at least one visual processing model, determine an initial visual processing model adapted to the target scene; Based on the scene input data of the target scene, the initial visual processing model is trained to obtain the target visual processing model; The model request includes model specification parameters; determining the corresponding target visual processing model from at least one visual processing model based on the model request includes: Based on the model specification parameters, the corresponding target visual processing model is searched from the model library, wherein the model library stores multiple visual processing models with different model specification parameters.

13. The method according to claim 11, further comprising, after determining the corresponding target visual processing model from at least one visual processing model based on the model request: The target visual processing model is deployed, and a visual processing interface is constructed based on the target visual processing model so that the smart terminal can schedule the target visual processing model to perform visual tasks.

14. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1-9 or any one of claims 11-13.

15. A smart terminal with a camera, the smart terminal further comprising a voice acquisition device, a visual task processing system, and a player; The camera is configured to capture at least one frame of image and write it into a reference image set; The voice acquisition device is configured to acquire task voice and generate target visual task; The visual task processing system is configured to acquire task parameters of the target visual task, wherein the task parameters include task information of the target visual task and a corresponding reference time range, the reference time range being the time range preceding the target visual task; select a target image sequence from a reference image set according to the reference time range; and obtain the target processing result of the target visual task through a target visual processing model based on the target image sequence and the task information. The player is configured to play the target processing results of the target visual task.

16. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the method of any one of claims 1-9 or any one of claims 11-13.

17. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method of any one of claims 1-9 or any one of claims 11-13.