An image data transmission method and apparatus

By training models and parameters on cloud servers and performing differential encoding on camera devices, the problems of bandwidth consumption and resource waste in static scenes are solved, and efficient, real-time video transmission is achieved.

CN122120410APending Publication Date: 2026-05-29HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2024-11-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing video encoding technologies suffer from severe bandwidth consumption and resource waste when processing static scenes, and are difficult to implement real-time video transmission on resource-constrained devices.

Method used

The model and parameters are trained by a cloud server, and the camera device performs differential encoding to transmit only the parts of the scene that change. The model and parameters are optimized by combining neural network technology to achieve efficient video transmission.

Benefits of technology

Significantly reduces bandwidth requirements, improves video transmission efficiency and real-time performance, reduces resource consumption, and adapts to the flexibility needs of different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122120410A_ABST
    Figure CN122120410A_ABST
Patent Text Reader

Abstract

The application provides an image data transmission method and device to reduce the transmission bandwidth when transmitting image data, and relates to the technical field of image processing. In the method, a cloud server receives at least one fourth image captured by a camera device. The cloud server trains a first model and a first parameter based on the at least one fourth image. The first model is used to indicate the environment captured by the camera device, and the first parameter is used to represent the change parameter of the environment captured by the camera device. The cloud server sends the first model and the first parameter. The cloud server receives encoded difference data, which is the difference data between a first image and a second image. Based on the above scheme, compared with the video encoding and transmission scheme in the related art, the above scheme can transmit only the changed part of the scene through the difference encoding technology, and significantly reduce the bandwidth requirement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image data transmission method and apparatus. Background Technology

[0002] With the rapid development of network technology and the widespread application of surveillance technology, the real-time transmission of high-resolution surveillance video has become an increasingly important topic. Especially in the field of wireless communication, the bandwidth consumption of video surveillance systems directly affects the effective utilization of network resources and the overall cost of the system. Traditional surveillance video transmission schemes usually rely on the real-time transmission of full-resolution video. While this method can ensure the integrity of the video, it is inefficient and wastes resources when processing large amounts of redundant and static scene data.

[0003] In many surveillance scenarios, such as parking lots, streets, and building interiors, the environment often exhibits prolonged static or semi-static characteristics. This means that the video stream contains a large number of redundant frames, with minimal or almost non-existent differences between them. However, existing video coding technologies, such as Advanced Video Coding (H.264 / AVC) and High Efficiency Video Coding (H.265 / HEVC), are primarily designed for dynamic video data and do not adequately consider optimization for static scenes. Even in static scenes, the continuous output of full-resolution video consumes significant bandwidth resources, a problem particularly pronounced in high-resolution video surveillance systems. Summary of the Invention

[0004] This application provides an image data transmission method and apparatus to reduce the transmission bandwidth when transmitting image data.

[0005] Firstly, an image data transmission method is provided, which can be executed by a cloud server. In this method, the cloud server receives at least one fourth image captured by a camera device. Based on the at least one fourth image, the cloud server trains a first model and first parameters. The first model indicates the environment captured by the camera device, and the first parameters characterize changes in the environment captured by the camera device. The cloud server sends first information, which includes the first model and the first parameters. The cloud server receives encoded differential data, which is the difference between a first image and a second image. The first image is captured by the camera device after capturing at least one fourth image, and the second image is obtained based on the first model and the first parameters.

[0006] Based on the above solution, compared to video encoding and transmission schemes in related technologies, this solution can significantly reduce bandwidth requirements by transmitting only the changing parts of the scene through differential coding technology. Under resource-constrained conditions, such as with mobile monitoring devices, traditional video transmission schemes may struggle to achieve real-time video transmission. The above solution, through the interaction between a cloud server and the camera device, and using a first model and first parameters, adapts to the real-time and flexibility requirements of different scenarios.

[0007] In one possible implementation, after receiving the encoded differential data, the cloud server decodes the encoded differential data to obtain more differential data. The cloud server then performs at least one of the following operations: determines a second image, and determines a third image based on the differential data and the second image. Alternatively, it optimizes a first parameter based on the differential data.

[0008] Based on the above scheme, the cloud server can render a second image using the first model and first parameters. Furthermore, by using differential data and the second image, it can reconstruct the real image captured by the camera, thus achieving high-efficiency video transmission. In addition, the cloud server can optimize the first parameters based on the differential data, dynamically adjusting them to make the second image rendered using the first model and first parameters more closely resemble the real environment, reflecting dynamic changes in the environment and improving video transmission performance.

[0009] In one possible implementation, when the optimized first parameters fail to converge, the cloud server optimizes the first model. Based on the above scheme, the cloud server can optimize the first model based on images captured by the camera device, enabling the first model to reflect the real environment.

[0010] In one possible implementation, after optimizing the first model, the cloud server sends a second model, which is different from the first model and is obtained based on a third image.

[0011] In one possible implementation, the second model is also derived from the first model.

[0012] In one possible implementation, the second model is also determined based on at least one fourth image.

[0013] Based on the above solution, the cloud server uses neural network technology to build and optimize the environment model, dynamically adjust the changing parameters and environment model, and establish an efficient interaction mechanism between the cloud server and the camera device. This effectively overcomes the limitations of related technologies when processing static scenes, and achieves bandwidth optimization, computing resource saving, and real-time performance improvement in video data transmission.

[0014] Secondly, an image data transmission method is provided, which can be executed by a camera device. In this method, a first image is acquired. Differential data between the first image and a second image is encoded. The encoded differential data is then transmitted. The second image is obtained based on a first model and first parameters, whereby the first model indicates the environment captured by the camera device, and the first parameters characterize changes in the environment captured by the camera device.

[0015] In one possible implementation, first information is received, which includes a first model and first parameters.

[0016] In one possible implementation, after sending the encoded differential data, an optimized first parameter is received, which is determined based on the differential data.

[0017] In one possible implementation, at least one fourth image captured by the camera device is sent before acquiring the first image. This fourth image is used to obtain the first model and the first parameters.

[0018] In one possible implementation, a second model is received, which is different from the first model. The second model is obtained based on a third image, which is obtained based on the second image and the difference data.

[0019] In one possible implementation, the second model is also derived from the first model.

[0020] In one possible implementation, the second model is further determined based on at least one fourth image captured by the camera device. The at least one fourth image is used to obtain the first model and the first parameters.

[0021] Thirdly, an image data transmission device is provided, including a processing unit and a transceiver unit.

[0022] The transceiver unit is configured to receive at least one fourth image captured by the camera device. The processing unit is configured to train a first model and first parameters based on the at least one fourth image. The first model indicates the environment captured by the camera device, and the first parameters characterize the changing parameters of the environment captured by the camera device. The transceiver unit is also configured to transmit first information, which includes the first model and the first parameters. The transceiver unit is also configured to receive encoded differential data. The differential data is the difference between the first image and the second image. The first image is captured by the camera device after capturing at least one fourth image, and the second image is obtained based on the first model and the first parameters.

[0023] In one possible implementation, after receiving the encoded differential data, the processing unit is further configured to decode the encoded differential data to obtain differential data. The processing unit is also configured to perform at least one of the following operations: determine a second image; determine a third image based on the differential data and the second image; or optimize a first parameter based on the differential data.

[0024] In one possible implementation, the processing unit is also used to optimize the first model when the optimized first parameters do not converge.

[0025] In one possible implementation, the transceiver unit is also used to optimize the first model and then send a second model, which is different from the first model and is obtained based on a third image.

[0026] In one possible implementation, the second model is also derived from the first model.

[0027] In one possible implementation, the second model is also determined based on at least one fourth image.

[0028] Fourthly, an image data transmission device is provided, including a processing unit and a transceiver unit.

[0029] The processing unit is used to acquire a first image. The processing unit is also used to encode the differential data between the first image and the second image. The transceiver unit is used to transmit the encoded differential data. The second image is obtained based on a first model and first parameters. The first model indicates the environment captured by the camera device, and the first parameters characterize the changing parameters of the environment captured by the camera device.

[0030] In one possible implementation, the transceiver unit is also used to receive first information, which includes a first model and first parameters.

[0031] In one possible implementation, after sending the encoded differential data, the transceiver unit is also used to receive the optimized first parameter, which is determined based on the differential data.

[0032] In one possible implementation, before acquiring the first image, the transceiver unit is further configured to transmit at least one fourth image captured by the camera device. This at least one fourth image is used to obtain the first model and the first parameters.

[0033] In one possible implementation, the transceiver unit is also used to receive a second model, which is different from the first model. The second model is obtained based on a third image, which is obtained based on the second image and differential data.

[0034] In one possible implementation, the second model is also derived from the first model.

[0035] In one possible implementation, the second model is further determined based on at least one fourth image captured by the camera device. The at least one fourth image is used to obtain the first model and the first parameters.

[0036] Fifthly, this application provides a computer-readable storage medium storing a computer program or instructions that, when executed by a processor, cause an image data transmission apparatus to perform the method in any possible implementation of the first or second aspect described above.

[0037] Sixthly, this application provides a computer program product comprising a computer program or instructions that, when executed by a processor, cause an image data transmission apparatus to perform the method in any possible implementation of the first or second aspect described above.

[0038] In a seventh aspect, this application provides a chip including a processor coupled to a memory for executing a computer program or instructions stored in the memory, such that the chip implements the method in any possible implementation of the first or second aspect.

[0039] Eighthly, this application provides an image data transmission apparatus, including at least one processor and an interface. The at least one processor is configured to read instructions via the interface to execute methods as described in any possible implementation of the first or second aspect. Attached Figure Description

[0040] Figure 1 A schematic diagram of a system architecture provided for an embodiment of this application;

[0041] Figure 2 An exemplary flowchart of an image data transmission method provided in an embodiment of this application;

[0042] Figure 3A A schematic diagram of a scenario for an image data transmission method provided in an embodiment of this application;

[0043] Figure 3B A schematic diagram illustrating another image data transmission method provided in this application embodiment;

[0044] Figure 4 A schematic diagram of an apparatus provided in an embodiment of this application;

[0045] Figure 5 This is a schematic diagram of another device provided in an embodiment of this application. Detailed Implementation

[0046] For ease of description, the technical terms involved in the embodiments of this application will be explained and described below.

[0047] 1) High Efficiency Video Coding (H.265 / HEVC) is a high-efficiency video compression standard designed to achieve video transmission while maintaining high video quality under low bandwidth conditions. Compared to its predecessor, Advanced Video Coding (H.264 / AVC), H.265 / HEVC introduces the following techniques to improve compression efficiency:

[0048] • Coding units (CUs): HEVC divides the video frame into multiple coding units of varying sizes. Each unit can be independently predicted, transformed, and quantized, which provides finer coding granularity and greater coding flexibility.

[0049] • Prediction techniques: HEVC supports larger prediction block sizes and finer motion vector accuracy, while also providing more prediction modes, including multiple directions of intra-frame prediction and inter-frame prediction.

[0050] • Transformation and quantization: HEVC supports larger transform blocks, up to 64x64 in size, and uses adaptive transform and quantization techniques, which helps to remove redundant information from video frames more effectively.

[0051] • Entropy coding: HEVC employs more advanced entropy coding techniques, such as context-adaptive binary arithmetic coding (CABAC), to encode prediction information and residual data more efficiently.

[0052] The combination of these technologies allows H.265 / HEVC to provide a higher compression rate than H.264 / AVC without sacrificing video quality, thus enabling the transmission of higher-quality video streams under limited bandwidth conditions. Although the H.265 / HEVC standard has achieved significant improvements in video compression efficiency, it still has some limitations and shortcomings, especially when processing static scene videos, where the following problems arise:

[0053] • High encoding complexity: HEVC's encoding complexity is far higher than H.264 / AVC, mainly due to its finer-grained coding unit division and more complex prediction modes. The complex encoding process requires high processing power, which can become a bottleneck for resource-constrained edge devices, such as surveillance cameras.

[0054] • Insufficient optimization for static scenes: Although HEVC offers various prediction modes, their efficiency is not fully utilized when processing static scenes. For long-term static monitoring footage, even with HEVC, a large amount of full-frame video data is still transmitted, resulting in wasted bandwidth.

[0055] • Real-time performance issues: Due to the complexity of the encoding process, HEVC may face increased latency when transmitting real-time video at high resolutions. This could cause inconvenience in practical applications for scenarios requiring real-time monitoring.

[0056] The aforementioned shortcomings indicate that although the H.265 / HEVC standard has achieved significant breakthroughs in video compression, more efficient technical solutions are still needed when processing static scene videos and resource-constrained edge devices to overcome the resource waste and real-time challenges in current video transmission. While the H.265 / HEVC standard has improved video compression efficiency, its encoding process is highly complex, especially in processing static scene videos, where it fails to fully utilize scene characteristics for optimization. This standard is primarily designed for dynamic video data and fails to effectively distinguish and process static and dynamic content, resulting in the transmission of a large amount of redundant data in static scenes, increasing bandwidth consumption and computational costs. Furthermore, due to the complexity of the encoding process, HEVC may face real-time challenges in high-resolution video surveillance environments, affecting the timeliness of video transmission and scene response speed.

[0057] In conclusion, while the H.265 / HEVC video compression standard boasts high compression efficiency, it exhibits limitations in handling static monitoring scenarios and meeting real-time requirements, leaving room for further innovation and technological development in the field of video coding.

[0058] It should be noted that the "static scene" involved in the embodiments of this application is relative. The "static scene" can be understood as a dynamic target in the environment, such as a small number of people and vehicles, or as an environment with little change and no change for a long time.

[0059] 2) Full-resolution real-time video transmission technology is a commonly used data transmission method in surveillance video systems. Its core principle is to ensure the real-time nature and clarity of the video, enabling real-time, delay-free observation of the monitored scene. This technical solution mainly achieves video data acquisition, processing, and transmission through the following steps:

[0060] 1. Video capture: The surveillance camera continuously captures video footage at the highest supported resolution and frame rate.

[0061] 2. Data Encoding: The captured video data is encoded in real time, typically using widely supported video compression standards such as H.264 / AVC or H.265 / HEVC, to reduce the bandwidth requirements for data transmission.

[0062] 3. Wireless transmission: The encoded video data is transmitted in real time to the monitoring center or cloud server via wireless network technologies such as wireless fidelity (Wi-Fi) or 4th-generation (4G) / 5th-generation (5G) networks.

[0063] 4. Video Decoding and Playback: The receiving device, such as the monitor or mobile device in the monitoring center, decodes the received video data and plays the video in real time.

[0064] Although full-resolution real-time video transmission technology excels in real-time monitoring and high-definition display, it has several technical limitations and shortcomings, mainly reflected in:

[0065] • Huge bandwidth consumption: Full-resolution video data is massive, and even with efficient encoding technologies like H.265 / HEVC compression, it still requires significant network bandwidth. In wireless network environments, especially Wi-Fi and 4G / 5G networks, bandwidth limitations during peak hours can become a bottleneck for video backhaul, leading to transmission delays or quality degradation.

[0066] High cost: To meet the requirements of real-time transmission of full-resolution video, the monitoring system needs to deploy high-end network infrastructure and high-performance encoding and decoding equipment. This not only increases the initial investment in system hardware and network, but also raises daily operation and maintenance costs.

[0067] • Low utilization of computing resources: When processing video data, the full-resolution real-time video transmission technology fails to fully consider the dynamic changes in the scene. Whether in the encoding or decoding process, it uses the same data processing strategy for static or minimally changing monitoring footage as it does for dynamic scenes, leading to a waste of computing resources.

[0068] • Difficulty adapting to resource-constrained devices: In resource-constrained environments, such as battery-powered mobile monitoring devices or edge computing nodes, real-time full-resolution video transmission may not be possible due to insufficient computing and storage resources, limiting the application scenarios and flexibility of the monitoring system. As resolution increases, the bandwidth required for video also increases rapidly.

[0069] The above points demonstrate that while full-resolution real-time video transmission technology can meet the demands of high definition and real-time monitoring, its limitations become increasingly apparent in scenarios with limited network resources, cost sensitivity, limited computing resources, and the need for optimization of static scenes, posing challenges to the optimization and upgrading of video surveillance systems. In summary, full-resolution real-time video transmission technology has significant advantages in ensuring image quality and real-time performance, but it also has obvious limitations in terms of bandwidth consumption, cost, computing resource utilization, and adaptability to resource-constrained devices.

[0070] The specific embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0071] like Figure 1 The diagram shown is a schematic of a system architecture applicable to an embodiment of this application. Figure 1 The system architecture shown includes a camera device 100, a cloud server 200, and a communication network. The camera device 100 is used to acquire images, such as a camera or monitoring device. The cloud server 200 can be used for image processing, model-in-priority learning, etc. The camera device 100 and the cloud server 200 can communicate via the communication network. This communication network can be a 4G network, a 5G network, a 6th-generation (6G) mobile communication network, or a future communication network. Alternatively, the communication network can be a short-range communication network, such as satellite imagery, Bluetooth, or Wi-Fi.

[0072] With the rapid development of network technology and the widespread application of surveillance technology, the real-time transmission of high-resolution surveillance video has become an increasingly important topic. Especially in the field of wireless communication, the bandwidth consumption of video surveillance systems directly affects the effective utilization of network resources and the overall cost of the system. Traditional surveillance video transmission schemes usually rely on the real-time transmission of full-resolution video. While this method can ensure the integrity of the video, it is inefficient and wastes resources when processing large amounts of redundant and static scene data.

[0073] In many surveillance scenarios, such as parking lots, streets, and building interiors, the environment often exhibits long-term static or semi-static characteristics. This means that there are a large number of redundant frames in the video stream, with differences between these frames that may be minimal or almost non-existent. However, existing video coding technologies, such as H.264 / AVC and H.265 / HEVC, are primarily designed for dynamic video data and do not adequately consider optimization for static scenes. Even in static scenes, continuous output of full-resolution video consumes significant bandwidth resources, a problem particularly pronounced in high-resolution video surveillance systems.

[0074] Traditional video coding methods often employ techniques such as inter-frame prediction and intra-frame prediction to compress the entire video frame. However, as the resolution of surveillance cameras continues to increase, the complexity of video coding also increases, leading to longer encoding times and greater computational resource consumption. Furthermore, to ensure real-time video transmission and image quality, at high resolutions, coding algorithms often need to find a balance between compression ratio and image quality, which further limits the efficiency of data compression. Pursuing higher compression ratios while sacrificing image quality directly impacts the practicality of the surveillance system, reducing the analytical and recognition value of the video data.

[0075] Furthermore, current video encoding technologies fail to effectively distinguish between static and dynamic content in videos. When processing static scenes, the encoding process still requires processing the entire frame, failing to optimize for areas with minimal change. This limitation results in the transmission of large amounts of data even when the video content remains almost unchanged, leading to a waste of network resources. There is also a lack of intelligent analysis and processing of video content.

[0076] In view of this, embodiments of this application provide an image data transmission method. In this method, a cloud server can receive at least one fourth image captured by a camera device, and train a first model and first parameters based on the at least one fourth image. The cloud server can send the first model and first parameters to the camera device. The camera device can capture a first image and determine a second image based on the first model and first parameters. The camera device can encode the differential data between the first image and the second image and send the encoded differential data to the cloud server. In the above method, the first model can indicate the environment captured by the camera device, and the first parameters can characterize the changing parameters of the environment captured by the camera device.

[0077] Based on the above solution, compared to the video encoding and transmission solutions introduced earlier, this solution can transmit only the changing parts of the scene through differential coding technology, significantly reducing bandwidth requirements. Under resource-constrained conditions, such as mobile monitoring devices, traditional video transmission solutions may struggle to achieve real-time video transmission. The above solution, through the interaction between the cloud server and the camera device, and using a first model and first parameters, adapts to the real-time and flexibility requirements of different scenarios.

[0078] Furthermore, real-time transmission and processing of full-resolution video poses a challenge on devices with limited battery power or computing capabilities. The above solution trains the first model and first parameters on a cloud server, and the camera device performs differential encoding, which reduces the burden on resource-constrained devices and improves the applicability and popularity of the solution.

[0079] In summary, the technical solution provided in this application, through the application of intelligent differential coding and neural network technology, addresses the shortcomings of related technologies in terms of bandwidth, computing resources, real-time performance, video quality, equipment adaptability, and cost control, offering a more efficient, intelligent, and cost-effective video transmission solution. The technical solution provided in this application significantly improves the performance of video surveillance systems while reducing system maintenance and operating costs.

[0080] See Figure 2 The following is an exemplary flowchart of an image data transmission method provided in an embodiment of this application, which may include the following steps:

[0081] S201: The cloud server receives at least one fourth image captured by the camera device.

[0082] For example, the camera device can take a picture to obtain at least one fourth image. The camera device can then send the obtained fourth image to a cloud server.

[0083] S202: The cloud server trains the first model and first parameters based on at least one fourth image.

[0084] Figure 2 In the illustrated embodiment, the first model can indicate the environment being filmed by the camera device, and the first parameter can be used to characterize the changing parameters of the environment being filmed by the camera device. For example, the first parameter can characterize the lighting conditions being filmed by the camera device, such as weather, including rainy days, sunny days, foggy days, snowy days, etc., or time, including daytime, nighttime, etc.

[0085] In one possible implementation, the cloud server can use neural radiancefield (NeRF) technology to train a first model and first parameters. NeRF is a 3D scene reconstruction and rendering method that trains a neural network to represent the radiance field of a scene, thereby generating highly realistic images based on different viewpoints and lighting conditions. NeRF is trained through a series of images to ultimately learn a continuous 3D scene representation, capable of rendering scene images under arbitrary viewpoints and lighting conditions. For example, the cloud server can fit a model (first model) of the environment captured by the camera device based on at least one fourth image, i.e., learn a continuous 3D representation of the environment.

[0086] It should be noted that cloud servers can also use other technologies to train the first model and first parameters, such as 3D Gaussian Splatting technology, deep neural networks, etc.

[0087] S203: The cloud server sends the first information to the camera device.

[0088] Correspondingly, the camera device receives the first information from the cloud server.

[0089] The first information may include the first model and the first parameters.

[0090] S204: The camera device acquires the first image.

[0091] For example, a camera device can take a picture to obtain a first image, such as... Figure 3A As shown.

[0092] S205: The camera device encodes the differential data of the first image and the second image.

[0093] For example, see Figure 3B As shown, the camera device can determine the second image based on the first model and the first parameters. For example, the camera device can render the second image based on the first model and the first parameters. In S205, the camera device can identify the first image and the second image and determine the difference data between the first image and the second image. The camera device can encode the difference data.

[0094] It should be noted that the camera device may use encoding methods such as H.265 or H.264, or other encoding methods, when encoding differential data; this application does not impose any specific limitations.

[0095] S206: The camera device sends encoded differential data to the cloud server.

[0096] Correspondingly, the cloud server receives encoded differential data from the camera device.

[0097] In one possible implementation, the cloud server can decode the encoded differential data to obtain differential data. In another possibility, the cloud server can determine a third image based on the differential data and the second image. For example, the cloud server can render a second image based on a first model and first parameters, and then determine the third image based on the differential data and the second image. Exemplarily, the cloud server can send the third image to a display device for display.

[0098] Based on the above scheme, the camera device can send differential data to the cloud server, which can effectively reduce bandwidth consumption and avoid repeated encoding of static environments, thereby improving transmission efficiency and utilization of computing resources. The cloud server can render a second image using the first model and first parameters, and then reconstruct the real image captured by the camera device using the differential data and the second image, achieving highly efficient video transmission.

[0099] In another possible scenario, the cloud server can optimize the first parameter based on differential data. For example, the cloud server can optimize the first parameter based on differential data using NeRF technology. Optionally, the cloud server can send the optimized first parameter to the camera device. In this way, the camera device can determine the second image based on the first model and the optimized first parameter.

[0100] Based on the above scheme, the cloud server can dynamically optimize the first parameter, making the second image rendered by the first model and the first parameter closer to the real environment, able to reflect the dynamic changes of the environment, and improve the performance of video transmission.

[0101] In some embodiments, the first parameter may fail to converge when the cloud server optimizes it. For example, if the cloud server did not encounter the lighting conditions in the difference data during the training of the first model and the first parameter, the first parameter may fail to converge when the cloud server optimizes it. In this case, the cloud server can optimize the first model.

[0102] For example, a cloud server can optimize a first model and first parameters based on a third image to obtain a second model and second parameters. Alternatively, a cloud server can train a second model and second parameters based on at least one fourth image and a third image.

[0103] Optionally, the cloud server can send the second model and the second parameters to the camera device. In this way, the camera device can determine the second image based on the second model and the second parameters.

[0104] Based on the above solution, the cloud server can optimize the first model and first parameters based on images captured by the camera device, enabling the first model and first parameters to reflect the real environment. The solution provided in this application, through neural network technology, constructs and optimizes the environment model, dynamically adjusts changing parameters, implements differential coding strategies, and establishes an efficient interaction mechanism between the cloud server and the camera device. This effectively overcomes the limitations of related technologies in processing static scenes, achieving bandwidth optimization, computational resource conservation, and improved real-time performance for video data transmission.

[0105] The solutions provided in this application can be applied to high-definition public safety monitoring scenarios. In urban public safety monitoring systems, high-resolution cameras are deployed throughout key areas. Figure 2The technical solution shown allows the first model and first parameters to construct the environment of a key area. The camera can upload differential data to the monitoring center, which can then reconstruct the true image based on the differential data, the first model, and the first parameters. Therefore, it significantly reduces bandwidth consumption, ensuring that surveillance video can be transmitted to the monitoring center in real-time and efficiently with limited network resources, while maintaining high-quality video footage and improving the speed of security incident identification and response.

[0106] The above solution can also be applied to traffic flow monitoring and analysis scenarios. Surveillance cameras in intelligent transportation systems are used to monitor road conditions and vehicle movement. Through... Figure 2 The technical solution shown allows the first model and first parameters to construct an environment representing road conditions. The camera can upload differential data on vehicle movement to the monitoring center, enabling the center to determine dynamic changes in vehicles, effectively reducing network resource consumption and supporting large-scale traffic monitoring network deployment.

[0107] Furthermore, the above solution can also be applied to intelligent retail analytics scenarios. Surveillance cameras deployed in retail stores are used to analyze customer traffic and the status of merchandise on shelves. Figure 2 The technical solution shown, with its first model and first parameters, can construct a shelf merchandise environment. The camera can send changes in customer behavior and merchandise placement to the monitoring center, enabling real-time analysis of customer behavior and merchandise management, and improving the operational efficiency of retail operations.

[0108] The above scenarios are merely illustrative examples. The technical solutions provided in this application can also be applied to other scenarios, such as resource-constrained scenarios with limited computing resources or limited transmission resources.

[0109] Figure 4 A schematic diagram of a device 400 is shown. This device 400 can be the aforementioned camera device or cloud server, or a device built into the camera device or cloud server, capable of implementing the image data transmission method provided in this application embodiment. The device 400 can be a hardware structure, a software module, or a hardware structure plus a software module. The device 400 can be implemented by a chip system. In this application embodiment, the chip system can be composed of chips or may include chips and other discrete components.

[0110] The device 400 may include a transceiver unit 401 and a processing unit 402.

[0111] For example, transceiver unit 401 is used to receive at least one fourth image captured by the camera device. Processing unit 402 is used to train a first model and first parameters based on the at least one fourth image. The first model is used to indicate the environment captured by the camera device, and the first parameters are used to characterize the changing parameters of the environment captured by the camera device. Transceiver unit 401 is also used to transmit first information, which includes the first model and the first parameters. Transceiver unit 401 is also used to receive encoded differential data. The differential data is the difference between the first image and the second image. The first image is captured by the camera device after capturing at least one fourth image, and the second image is obtained based on the first model and the first parameters.

[0112] For example, processing unit 402 is used to acquire a first image. Processing unit 402 is also used to encode the differential data between the first image and the second image. Transceiver unit 401 is used to transmit the encoded differential data. The second image is obtained based on a first model and first parameters. The first model indicates the environment captured by the camera device, and the first parameters characterize the changing parameters of the environment captured by the camera device.

[0113] For the specific execution process of the transceiver unit 401 and the processing unit 402, please refer to the description in the above method embodiments. The module division in this application embodiment is illustrative and only represents a logical functional division. In actual implementation, there may be other division methods. Furthermore, the functional modules in the various embodiments of this application can be integrated into a processor, exist as separate physical entities, or have two or more modules integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0114] Figure 5 This is a schematic diagram of the hardware structure of a device 500 provided in an embodiment of this application. Figure 5 The illustrated device 500 includes a memory 501, a processor 502, a communication interface 503, and a bus 504. The memory 501, processor 502, and communication interface 503 are interconnected via the bus 504.

[0115] The memory 501 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 501 may store a program. When the program stored in the memory 501 is executed by the processor 502, the processor 502 and the communication interface 503 are used to execute the various steps of the image data transmission method of the embodiments of this application.

[0116] The processor 502 may be a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), graphics processing unit (GPU), or one or more integrated circuits, used to execute related programs to achieve the functions required by the units or modules in the camera device or cloud server of this application embodiment, or to execute the image data transmission method provided in this application embodiment.

[0117] The processor 502 can also be an integrated circuit chip with signal processing capabilities. In implementation, each step of the image data transmission method in the above embodiments can be completed by the integrated logic circuits in the hardware of the processor 502 or by instructions in software form. The processor 502 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the memory 501. The processor 502 reads the information in the memory 501 and, in conjunction with its hardware, completes the image data transmission method of this application embodiment.

[0118] The communication interface 503 uses transceiver devices, such as, but not limited to, transceivers, to enable communication between the device 500 and other devices or communication networks. For example, when the device 500 is a cloud server, it can receive images from a camera device through the communication interface 503.

[0119] Bus 504 may include a pathway for transmitting information between various components of device 500 (e.g., memory 501, processor 502, communication interface 503).

[0120] It should be understood that the transceiver unit 401 in device 400 is equivalent to the communication interface 503 in device 500, and the processing unit 402 can be equivalent to the processor 502.

[0121] It should be noted that, although Figure 5 The illustrated device 500 only shows the memory, processor, and communication interface. However, those skilled in the art should understand that in specific implementations, device 500 may also include other devices necessary for normal operation. Furthermore, depending on specific needs, those skilled in the art should understand that device 500 may also include hardware devices for implementing other additional functions. Moreover, those skilled in the art should understand that device 500 may only include the devices necessary for implementing the embodiments of this application, and may not necessarily include... Figure 5 All the devices shown.

[0122] This application also provides a computer-readable medium storing program code for execution by a device, the program code including the image data transmission method described in the foregoing embodiments.

[0123] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to execute the image data transmission method described in the foregoing embodiments.

[0124] This application also provides a chip, which includes a processor and a data interface. The processor reads instructions stored in a memory through the data interface and executes the image data transmission method described in the foregoing embodiments.

[0125] In one possible design, the chip may further include a memory storing instructions, and the processor is used to execute the instructions stored in the memory. When the instructions are executed, the processor is used to execute the image data transmission method described in the foregoing embodiments.

[0126] This application also provides an electronic device, which includes a processing apparatus for performing the image data transmission method described in the foregoing embodiments.

[0127] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0128] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0129] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0130] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0131] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0132] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0133] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image data transmission method, characterized in that, Applied to camera devices, including: Get the first image; The difference data between the first image and the second image is encoded; Send the encoded differential data; The second image is obtained based on a first model and first parameters. The first model is used to indicate the environment captured by the camera device, and the first parameters are used to characterize the changing parameters of the environment captured by the camera device.

2. The method according to claim 1, characterized in that, Also includes: Receive first information, which includes the first model and the first parameters.

3. The method according to claim 1 or 2, characterized in that, After transmitting the encoded differential data, the method further includes: The first parameter for optimization is received, which is determined based on the differential data.

4. The method according to any one of claims 1 to 3, characterized in that, Before acquiring the first image, the process also includes: Send at least one fourth image captured by the camera device; wherein the at least one fourth image is used to obtain the first model and the first parameters.

5. The method according to claim 4, characterized in that, Also includes: Receive a second model, which is different from the first model. The second model is obtained based on a third image, which is obtained based on the second image and the difference data.

6. The method according to claim 4, characterized in that, The second model is also derived from the first model.

7. The method according to claim 5, characterized in that, The second model is further determined based on at least one fourth image captured by the camera device; wherein the at least one fourth image is used to obtain the first model and the first parameter.

8. An image data transmission method, characterized in that, include: Receive at least one fourth image captured by the camera device; Based on the at least one fourth image, a first model and first parameters are trained to obtain the first model; The first model is used to indicate the environment captured by the camera device, and the first parameter is used to characterize the changing parameters of the environment captured by the camera device; Send first information, the first information including the first model and the first parameter; Receive the encoded differential data; The difference data is the difference data between the first image and the second image; wherein the first image is obtained by the camera device after capturing the at least one fourth image, and the second image is obtained based on the first model and the first parameters.

9. The method according to claim 8, characterized in that, After receiving the encoded differential data, the method further includes: The encoded differential data is decoded to obtain differential data; Perform at least one of the following operations: Determine the second image, and based on the differential data and the second image, determine the third image; or, The first parameter is optimized based on the differential data.

10. The method according to claim 9, characterized in that, Also includes: If the optimized first parameter does not converge, optimize the first model.

11. The method according to claim 10, characterized in that, After optimizing the first model, the method further includes: Send a second model, which is different from the first model, and is obtained based on the third image.

12. The method according to claim 11, characterized in that, The second model is also derived from the first model.

13. The method according to claim 11, characterized in that, The second model is also determined based on the at least one fourth image.

14. An apparatus, characterized in that, The device includes at least one processor coupled to at least one memory; the at least one processor is configured to execute a computer program or instructions stored in the at least one memory to cause the device to perform the method as claimed in any one of claims 1 to 6, or to cause the device to perform the method as claimed in any one of claims 7 to 13.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program or instructions that, when read and executed by a computer, cause the computer to perform the method as described in any one of claims 1 to 6, or cause the computer to perform the method as described in any one of claims 7 to 13.

16. A computer program product, characterized in that, When the computer program product is run on the processor, it causes the detection system to perform the method as described in any one of claims 1 to 6, or causes the detection system to perform the method as described in any one of claims 7 to 13.