Content uploading method, wearable device, and storage medium

By extracting text information from image data in wearable devices and selecting the upload method based on quality assessment, the high power consumption problem of wearable devices is solved, achieving a balance between improved battery life and high-quality interactive response.

CN122513591APending Publication Date: 2026-08-04GOERTEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GOERTEK INC
Filing Date
2026-04-13
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Wearable devices suffer from excessive power consumption in image data transmission, especially in unstable network environments and when images contain a large amount of redundant data, making it difficult to guarantee the device's battery life.

Method used

After receiving an interactive instruction, the system extracts text information from the image data and evaluates its quality. Based on decision rules, it selects whether to upload text or images to reduce redundant data transmission.

Benefits of technology

It effectively reduces unnecessary power consumption during data transmission, extends the device's battery life, and reduces the risk of privacy leaks while ensuring accurate execution of interactive commands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122513591A_ABST
    Figure CN122513591A_ABST
Patent Text Reader

Abstract

The application discloses a content uploading method, a wearable device and a storage medium, and relates to the technical field of wearable devices. The method comprises the following steps: after receiving an interaction instruction, acquiring collected image data; extracting text information in the image data to obtain text data, and selecting at least one first index for evaluating the quality of the text data; acquiring or setting a decision rule corresponding to the first index, and deciding an uploading mode according to the decision rule and the first index, wherein the uploading mode is uploading text or uploading an image; and uploading corresponding data to a target device according to the uploading mode, so that the target device executes the interaction instruction based on the received data. The application reduces the invalid image transmission process and reduces the data transmission power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a content uploading method, wearable device, and storage medium. Background Technology

[0002] With the rapid development of mobile terminals and IoT technologies, interaction methods based on image acquisition and content recognition have been widely applied to various wearable devices. Users capture images, and the device or cloud extracts and processes the text information in the images to respond to the user's interactive commands.

[0003] In current interactive scenarios, after a wearable device captures an image containing text, it uploads the complete image data to the target device, such as a mobile terminal or the cloud. The target device then analyzes and recognizes the image to execute user commands.

[0004] However, traditional content upload methods face the common problem of excessive power consumption in practical applications of wearable devices. Due to the inherently limited battery life of wearable devices, coupled with unstable network environments and the presence of large amounts of redundant data in images, ineffective power consumption during data transmission remains high, making it difficult to effectively guarantee a balanced device battery life. Therefore, reducing ineffective image transmission processes and lowering data transmission power consumption have become urgent technical problems to be solved. Summary of the Invention

[0005] The main purpose of this application is to provide a content uploading method, a wearable device, and a storage medium, aiming to solve the technical problems of how to reduce invalid image transmission processes and reduce data transmission power consumption.

[0006] To achieve the above objectives, this application provides a content uploading method, which includes: After receiving the interactive command, acquire the collected image data; Text information is extracted from the image data to obtain text data, and at least one first indicator is selected to evaluate the quality of the text data; Obtain or set the decision rule corresponding to the first indicator, and determine the upload method according to the decision rule and the first indicator, wherein the upload method is to upload text or upload an image; The corresponding data is uploaded to the target device according to the upload method described above, so that the target device can execute the interaction command based on the received data.

[0007] In addition, to achieve the above objectives, this application also provides a wearable device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the method described above.

[0008] In addition, to achieve the above objectives, this application also provides a readable storage medium, which is a computer-readable storage medium storing a computer program that is executed by a processor to implement the steps of the above method.

[0009] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described above.

[0010] One or more technical solutions proposed in this application have at least the following technical effects: This application, upon receiving an interactive command, acquires the collected image data but does not directly upload the complete image. Instead, it first extracts text information from the image data to obtain text data, selects at least one first indicator for evaluating the quality of the text data, and then acquires or sets the decision rule corresponding to the first indicator. Based on the decision rule and the first indicator, it determines the upload method (upload text or upload an image). This effectively avoids the drawbacks of traditional methods that always upload complete images, leading to redundant data transmission. By selecting the first indicator and combining it with the decision rule to flexibly determine the upload method, it allows for adaptive selection of appropriate upload content based on the text data quality. This avoids invalid transmission of uploading complete images regardless of the actual text quality, reduces the amount of redundant data transmission, and thus reduces the device's power consumption during data transmission. Therefore, this technical solution, by establishing a mechanism for text quality evaluation and adaptive upload method decision-making on the device side, effectively reduces invalid image transmission processes. While ensuring accurate execution of interactive commands, it reduces data transmission power consumption, thereby helping to extend the device's battery life and achieving a balance between battery life and interactive response quality. Attached Figure Description

[0011] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0012] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 A flowchart illustrating the first embodiment of the content upload method of this application; Figure 2 A schematic diagram illustrating the text content merging process involved in one embodiment of the content uploading method of this application; Figure 3A schematic diagram of the smart glasses system architecture involved in one embodiment of the content uploading method of this application; Figure 4 A schematic diagram of the threshold decision-making process involved in one embodiment of the content upload method of this application; Figure 5 A schematic diagram of the model decision-making process involved in one embodiment of the content upload method of this application; Figure 6 This application's content upload method is illustrated in an embodiment of a data acquisition and model training process diagram. Figure 7 This is a schematic diagram of the hardware operating environment involved in the content uploading method apparatus in the embodiments of this application.

[0014] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0015] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] Existing wearable devices such as smart glasses often upload captured images directly to a mobile phone or cloud for processing when performing tasks such as "recognition, translation, summarization, and information extraction." While this method is simple to implement, uploading entire images leads to unnecessary bandwidth and energy consumption in scenarios heavily reliant on text. Furthermore, image transmission introduces additional latency and increases privacy risks. Current solutions largely rely on fixed rules or manual selection, making it difficult to consistently determine "whether image uploading is necessary" under different content and tasks.

[0017] Based on this, the main solution of this application is as follows: after receiving the interaction command, acquire the collected image data; extract the text information from the image data to obtain text data, and select at least one first indicator for evaluating the quality of the text data; acquire or set the decision rule corresponding to the first indicator, and decide the upload method according to the decision rule and the first indicator, wherein the upload method is to upload text or upload an image; upload the corresponding data to the target device according to the upload method, so that the target device can execute the interaction command based on the received data.

[0018] This application, after acquiring image data, does not directly upload the complete image. Instead, it first extracts text information from the image to obtain text data, selects a first indicator to evaluate the quality of the text data, and then decides on the upload method (upload text or upload image) based on this first indicator. This effectively avoids the drawbacks of traditional methods that always upload complete images, leading to redundant data transmission. By calculating the target indicator and deciding on the upload method, it allows for flexible selection of appropriate upload content based on the quality of the text data, avoiding the invalid transmission of uploading complete images regardless of the actual data situation, reducing the amount of redundant data transmission, thereby reducing the unnecessary power consumption of wearable devices during data transmission, and also reducing the risk of privacy leakage.

[0019] It should be noted that the execution subject of each embodiment of the content uploading method of this application can be a wearable device capable of realizing the above functions, such as smart glasses, AR (Augmented Reality) headsets, VR (Virtual Reality) headsets, etc. The embodiments of the content uploading method of this application do not impose specific limitations on this. The following describes the embodiments of the content uploading method of this application with smart glasses as the execution subject.

[0020] Based on this, this application proposes a content uploading method according to a first embodiment. In the first embodiment, refer to... Figure 1 As shown, the content upload method includes the following steps S10~S40: Step S10: After receiving the interaction command, acquire the collected image data; The interaction command is triggered by the user through preset operations on the smart glasses. Preset operations include, but are not limited to, voice wake-up, touch operation, gesture command or physical button trigger. For example, the user can say "recognize the text in the picture", "extract key information", "summarize the main idea of ​​the article", etc., touch the touch area on the temple of the smart glasses, or make a preset gesture. All of these can be used as ways to trigger the interaction command.

[0021] The image data can be generated by the built-in image acquisition module (such as a miniature camera) of the smart glasses. The acquisition scenarios include, but are not limited to, paper documents, electronic screens, signs and other scenarios containing text information. The image data format can include common image formats such as JPG and PNG. After acquisition, it can be temporarily stored in the local cache of the smart glasses for subsequent steps.

[0022] Upon receiving an interaction command, the smart glasses can determine the acquisition strategy based on the command type. They can capture a single frame image from the camera's live stream, acquire multiple consecutive frames within a preset time window, or retrieve historical image data from the cache as the original image data to be uploaded.

[0023] Step S20: Extract text information from the image data to obtain text data, and select at least one first indicator for evaluating the quality of the text data; After acquiring image data, the smart glasses perform text recognition on the image. Specifically, this can be done by calling a text recognition algorithm, such as OCR (Optical Character Recognition), to extract text information from the image. The extracted text information is then arranged according to the original text order in the image to form structured text data. The text data can be in text format and temporarily stored in a local cache. For example, the following description and explanation of various embodiments of this application will use the extraction of text information from an image using an OCR recognition algorithm as an example.

[0024] Furthermore, before text recognition, the image can be preprocessed, such as grayscale conversion, binarization, tilt correction, and noise reduction, to improve the accuracy of subsequent recognition.

[0025] The smart glasses select at least one primary indicator to evaluate the quality of text data. This primary indicator quantifies the quality of the text data, such as its usability and recognition reliability. Specifically, the primary indicator may include, but is not limited to, one or more combinations of the following parameters: text clarity (e.g., the mean confidence score of each character), text proportion (the ratio of the text area to the total image area), text concentration (the degree of concentration of the text area in the image space), text completeness (e.g., the ratio of the number of recognized characters to the estimated number of characters), and text layout regularity (e.g., line spacing consistency).

[0026] Step S30: Obtain or set the decision rule corresponding to the first indicator, and determine the upload method according to the decision rule and the first indicator, wherein the upload method is to upload text or upload an image; The smart glasses acquire preset decision rules, or dynamically set decision rules based on the current interaction scenario. The decision rules are used to map the value of the first indicator to the upload method (uploading text data or uploading image data).

[0027] As an example, the decision rule could be a threshold-based decision rule: for instance, a quality threshold is pre-set. When the first indicator representing the quality of text data reaches or exceeds this threshold, it indicates that the locally extracted text data has high reliability and integrity, sufficient to support the target device in accurately executing interactive commands. In this case, the decision is to upload the text. Conversely, when the first indicator is below the threshold, it indicates that the text data may have recognition errors, information loss, or incomplete recognition. If such text data is directly uploaded, the target device may execute commands based on erroneous information, leading to operation failure or a degraded user experience. Therefore, the decision is to upload an image, allowing the target device to re-analyze and recognize the image.

[0028] In practical applications, this decision-making logic can be further refined, for example, by setting multi-level thresholds and using compromise solutions such as compressed image uploading when the quality is moderate. Preferably, the entire decision-making process can be completed locally on the smart glasses without needing to interact and negotiate with the target device, thereby effectively reducing decision latency and communication overhead.

[0029] As another example, the decision rule can be a model decision rule: the smart glasses, based on a pre-trained decision model, input the currently extracted first indicator (and / or other auxiliary features) into the model and directly obtain the upload method of the model output.

[0030] Through the above decision-making process, the smart glasses can adaptively choose between "uploading text" and "uploading images" to avoid redundant transmission caused by always uploading complete images.

[0031] Step S40: Upload the corresponding data to the target device according to the upload method, so that the target device can execute the interaction command based on the received data.

[0032] Based on the decision, the smart glasses can upload the corresponding data to the target device via their built-in communication modules (such as Wi-Fi, Bluetooth, cellular networks, etc.). For example, if the decision is to upload text, the smart glasses can package and send the extracted text data, which can also include metadata such as timestamps, recognition confidence levels, and text region coordinates to assist the target device in quickly parsing and responding.

[0033] If the decision is to upload an image, the acquired raw image data can be uploaded. Furthermore, while maintaining recognizability, appropriate compression or partial cropping can be performed to reduce the amount of data transmitted. After receiving the data, the target device executes the interaction command based on the received data. In an exemplary embodiment, after receiving the data, the target device inputs the received data and the user's question data into the VLM (Vision-Language Model), and outputs the interaction result to complete the response to the interaction command.

[0034] By employing the above methods, smart glasses prioritize text data with small transmission volumes while ensuring that the target device can accurately execute interactive commands. This effectively reduces the transmission of redundant data and lowers communication power consumption, making them particularly suitable for practical application scenarios where wearable devices have limited battery life and network environments fluctuate frequently.

[0035] Based on the first embodiment of the content uploading method of this application, in the second embodiment of the content uploading method of this application, the content that is the same as or similar to that in the first embodiment can be referred to the above description and will not be repeated hereafter. Based on this, the step of setting the decision rule corresponding to the first indicator includes: Step A10: Set the upper limit threshold and lower limit threshold of text quality corresponding to the first indicator, so as to compare with the first indicator and decide the upload method; Since the first indicator can include multiple parameters for evaluating the quality of text data, such as text clarity, text proportion, and text concentration, smart glasses can set corresponding upper and lower threshold values ​​for text quality for each first indicator.

[0036] It should be noted that the upper and lower threshold values ​​for text quality can be adaptively adjusted based on the task type corresponding to the interaction command. When the smart glasses receive an interaction command, they can simultaneously parse the task type information contained in the command. Different types of tasks have different requirements for text data quality. For example: When the task type is Extract, it usually requires accurate extraction of text content and has high requirements for recognition accuracy. A high lower limit threshold for text quality can be set for each indicator to ensure the reliability of the extraction results. At the same time, a more stringent combination judgment rule can be adopted, such as requiring all first indicators to be higher than the upper limit before it is judged as text that can be uploaded, and any indicator lower than the lower limit is judged as an image that needs to be uploaded.

[0037] When the task type is Translate, the translation quality is highly dependent on the accuracy of the source text. A high quality threshold can be set to avoid translation errors caused by recognition mistakes. Strict judgment rules similar to those for Extract can also be set.

[0038] When the task type is Summarize or Explain, this type of task focuses on the semantic understanding of the text content and has a certain tolerance for errors in the recognition of individual characters. The lower limit threshold of text quality can be appropriately relaxed, and relatively lenient combination judgment rules can be adopted. For example, text can be uploaded if the core indicators (such as confidence and completeness) are higher than the upper limit, and the requirements for secondary indicators can be appropriately relaxed.

[0039] When the task type is Key Info (field extraction), such as extracting key fields like invoice amount, ID number, and date, the accuracy of the key information is extremely important. Even a single character error can have serious consequences. For this type of task, strict text quality thresholds can be set, and a rigorous judgment rule can be adopted: "Text can only be uploaded if all primary indicators are above the upper limit; if any indicator is below the lower limit, an image must be uploaded."

[0040] When the task type is Verify, it usually involves comparing the identification results with existing information. The requirement for identification accuracy depends on the importance of the specific content to be verified. The threshold and judgment rules can be dynamically adjusted according to the criticality of the verification items.

[0041] When the task type is Draft, such as generating a draft or first draft based on the recognized content, the tolerance for recognition errors of individual characters is relatively high. Relatively lenient quality thresholds and judgment rules can be set, such as requiring the main indicators to be higher than the upper limit before uploading the text.

[0042] When the task type is Share / Save, this type of operation has high requirements for the integrity of the content, but has a high tolerance for errors in individual non-critical characters. The threshold can be adjusted according to the user's preset preferences, and the rule of uploading text can be adopted if most indicators meet the requirements.

[0043] When the task type is Search, the search results depend on the accuracy of the keywords. If the keywords are not identified correctly, the search results may deviate from the user's intent. Therefore, a higher lower limit threshold for text quality can be set, and the confidence level of the identification of key fields should be given special attention.

[0044] Meanwhile, since the first indicator may contain multiple parameters, corresponding upper limit combination judgment rules and lower limit combination judgment rules can be configured to integrate the threshold comparison results of various indicators into a final upload method decision. The upper limit combination judgment rules can include, but are not limited to, the following forms: One rule is "all criteria met," meaning that the text is considered acceptable only when all primary indicators are greater than or equal to their respective upper limits for text quality. This rule is suitable for task types with extremely high accuracy requirements, such as Explain, Verify, and Search.

[0045] Another rule is "core compliance," which pre-divides the first indicator into core indicators and auxiliary indicators. Only the core indicators are required to be greater than or equal to the corresponding upper limit threshold for text quality; auxiliary indicators are not mandatory. This rule is suitable for task types that have high requirements for specific quality dimensions but are more tolerant of other dimensions. For example, Translate focuses more on recognition confidence and completeness, and the requirements for text layout regularity can be appropriately relaxed.

[0046] Another rule is "majority compliance," which requires a preset number or proportion of primary indicators to reach or exceed the corresponding text quality upper limit threshold. For example, when there are five primary indicators, at least four must meet the criteria to be considered as meeting the text upload requirements. This rule is suitable for task types with relatively balanced quality requirements but allowing for a certain degree of fluctuation.

[0047] Another rule is "weighted compliance," which assigns weights to each primary indicator, calculates a weighted compliance rate, and determines that the text upload condition is met when the weighted overall score reaches a preset threshold. This rule allows for more refined quality assessment.

[0048] Similarly, the rules for determining lower bound combinations can take, but are not limited to, the following forms: One rule is "any threshold," meaning that if any of the primary indicators is less than or equal to its corresponding text quality threshold, the image must be uploaded. This rule is suitable for tasks with strict requirements for all quality dimensions, as a significant deficiency in any one indicator can lead to task failure.

[0049] Another rule is "core threshold," which means that an image is only required to be uploaded if any of the core metrics is less than or equal to its corresponding text quality lower limit threshold. This rule is suitable for task types with strict requirements for core quality dimensions.

[0050] Another rule is "majority bottoming out," which means that when the first indicator of the preset quantity or proportion is lower than or equal to the corresponding text quality lower limit threshold, it is determined that an image needs to be uploaded.

[0051] Another rule is "weighted bottoming out," which assigns weights to each primary indicator, calculates the weighted bottoming out rate, and determines that the conditions for uploading an image are met when the weighted comprehensive score is lower than a preset threshold.

[0052] Through the flexible configuration of threshold differentiation settings and combined judgment rules based on task type, smart glasses can adaptively adjust decision boundaries according to the actual needs of different application scenarios, and optimize energy consumption while ensuring the performance of tasks.

[0053] Alternatively, in step A20, a decision model is trained based on a preset training dataset to determine the upload method using the trained decision model, wherein the training dataset includes at least the first indicator.

[0054] As an alternative to threshold-based decision-making, smart glasses can employ a machine learning model-based decision-making approach. Specifically, a large number of training samples are pre-collected. Each training sample includes: the text recognition results of the sample image, the extracted primary indicators (such as text clarity, text proportion, text concentration, etc.), and the manually labeled optimal upload method (uploading text or uploading an image). These samples are then combined to form a training dataset.

[0055] The training dataset must at least contain the first metric as input features. Optionally, it may further include a second metric related to image data quality (such as image sharpness, contrast, edge density, etc.) and image data embedding vectors (such as image semantic features extracted through a pre-trained convolutional neural network). The uploaded annotation method serves as the supervision label.

[0056] The smart glasses utilize this training dataset to train the decision-making model offline. The decision-making model can employ classification models such as logistic regression, support vector machines, random forests, or lightweight neural networks (such as MobileNet and TinyML models). During training, the model learns a non-linear mapping relationship from input features (primary indicators, etc.) to the upload method decision.

[0057] After training, the decision model is deployed locally on the smart glasses. In actual use, the smart glasses extract the first indicator (and optional second indicator and image embedding vector) for the current interaction task in real time, and input it into the trained decision model. The model directly outputs the recommended upload method (upload text or upload image).

[0058] Through model-based decision-making, smart glasses can more accurately balance transmission efficiency and task execution quality, further enhancing the intelligence level of adaptive decision-making.

[0059] In one possible implementation, the decision rule is a threshold decision, and the step of determining the upload method based on the decision rule and the first indicator includes: Step B10: If the first indicator is greater than or equal to the upper limit threshold of text quality, then the upload method is determined to be text upload; The first indicator can include multiple evaluation parameters, such as text clarity, text proportion, and text concentration. The smart glasses set corresponding upper threshold values ​​for each of the first indicators. When the first indicator is a single item, it is directly compared with its upper threshold. When the first indicator has multiple items, the compliance status of each indicator is considered based on a preset combination judgment rule. For example, using a full compliance rule: all first indicators must be greater than or equal to their respective upper thresholds before text can be uploaded. Taking a text clarity upper limit of 0.85, a text proportion upper limit of 0.70, and a text concentration upper limit of 0.80 as an example, if the currently recognized clarity is 0.92, the proportion is 0.75, and the concentration is 0.85, all three meet the standards, then the text data is determined to be uploaded.

[0060] Based on the above combination of judgments, when the text data quality is sufficiently reliable, the smart glasses will choose to upload text data that is small in size and low in power consumption, thus ensuring the efficient execution of interactive commands.

[0061] Step B20: If the first indicator is less than or equal to the lower limit threshold of text quality, then the upload method is determined to be uploading an image.

[0062] The smart glasses set a minimum threshold for text quality for each primary indicator. When there are multiple primary indicators, a comprehensive judgment is made based on a combination of rules. For example, a veto rule is used: text clarity is set as the core indicator. If the clarity is less than or equal to its minimum threshold (e.g., 0.50), the image will be directly rejected regardless of whether other indicators meet the standards, because a low recognition confidence level would render the text content unusable.

[0063] The above rules ensure that when the quality of text data fails to meet the task requirements, the smart glasses automatically switch to uploading raw image data, which is then re-identified or analyzed by the target device with stronger processing capabilities, thereby avoiding command execution failure due to text errors.

[0064] In one possible implementation, the step of deciding the upload method based on the decision rule and the first indicator further includes: Step C10: If the first indicator is less than the upper limit threshold of text quality and greater than the lower limit threshold of text quality, then a second indicator for evaluating the quality of the image data is selected, wherein the second indicator includes one or more of the following: image sharpness, image contrast, image exposure, image edge density, image grayscale entropy, non-text information intensity, target object area ratio and target object saliency. When the evaluation result of the first indicator falls between the upper and lower thresholds of text quality—meaning neither the conditions for uploading text nor the conditions for uploading images are triggered—it indicates that the current condition is in a fuzzy "medium quality" range. In this case, relying solely on the first indicator is insufficient to make the optimal decision, and more multi-dimensional evaluation information needs to be introduced. Therefore, the smart glasses further calculate a second indicator, which is used to characterize the image data quality.

[0065] The second indicator may include, but is not limited to, one or more combinations of the following parameters: image sharpness, image contrast, image exposure, image edge density, image grayscale entropy, non-textual information intensity, target object area ratio, and target object saliency.

[0066] Step C20: Merge each of the second indicators to obtain the image transmission requirement score, wherein the second indicator indicates that the better the quality of the image data, the higher the image transmission requirement score; To reduce computational power consumption, smart glasses can preprocess the original image, such as downsampling it to a fixed size (e.g., 160×160 or 224×224 pixels) and converting it to grayscale, thus reducing subsequent computation. Simultaneously, a text mask can be constructed based on the set of text blocks obtained from OCR recognition. The bounding boxes of all text blocks are mapped onto the downsampled image and filled with 1s, while the remaining areas are filled with 0s, used to distinguish text regions from non-text regions later.

[0067] The second indicator may include, but is not limited to, one or more of the following: Image clarity This is used to reflect the overall sharpness and detail richness of an image. As an example, the Laplacian variance can be used as a proxy for sharpness, which is calculated by measuring the variance of the grayscale values ​​after the image has been filtered using the Laplacian operator: , among which, clip( The function (,0,1) is the cutoff function, which restricts the calculation result to the interval [0,1]. The preset resolution threshold, For the second-order Laplace differential operation, This is a variance calculation. The lower the variance, the blurrier the image. In this case, the local text recognition result may be unreliable, and it is more likely that the image will be uploaded for processing by the target device or the user will be prompted to retake the image.

[0068] Image exposure With contrast This is used to reflect the brightness level and the degree of difference between light and dark in an image. As an example, the average gray level of an image can be calculated as a brightness index, and the standard deviation of gray levels can be calculated as a contrast index. It can statistically analyze the ratio of overexposed areas (grayscale value greater than 240) and underexposed areas (grayscale value less than 15), and calculate based on this ratio. ,like ,in, The preset threshold is used to define the boundary of "close to pure black / pure white", and N is the total number of pixels in the image.

[0069] Image edge density This is used to reflect the richness of edge information in an image. As an example, the Sobel operator can be used to calculate the gradient magnitude of an image, and the number of pixels with gradient magnitudes higher than a preset threshold can be counted. The proportion of these pixels to the total number of pixels is then calculated as the edge density. ,in, Gx represents the Sobel gradient component of the image in the horizontal direction (X-axis), Gy represents the Sobel gradient component of the image in the vertical direction (Y-axis), # is a counting symbol representing the "number of pixels that meet the condition", and t is a preset gradient threshold. Higher edge density usually indicates that the image contains more structural information (such as charts, wireframes, textures, etc.). In this case, uploading only text may result in the loss of important layout information, making image uploads more suitable.

[0070] Image grayscale entropy (Also known as texture complexity), it reflects the complexity of the gray-level distribution of an image. As an example, the entropy value of an image's gray-level histogram (such as one with 64 or 256 gray levels) can be calculated, for instance. ,in, , k is the gray level index, ranging from 1 to K, K is the total number of gray levels, usually 64 or 256, and p(k) is the pixel percentage of the k-th gray level. A higher entropy value indicates richer image information, which may include charts, photos, complex backgrounds, etc.

[0071] Non-textual information intensity This parameter reflects the edge richness of non-text regions in an image and is a key indicator for distinguishing between images with only text and mixed text. As an example, it can be combined with a text mask to calculate edge density only for non-text regions (regions with a mask value of 0), count the number of pixels in non-text regions with gradient magnitudes exceeding a threshold, divide this number by the total number of pixels in the non-text region, and obtain the non-text proportion score. Based on this non-text proportion score, a positive determination is made. .For example, ,in, Mtext is the text mask, α is the lower threshold, representing "negligible non-text edge density", and β is the upper threshold, representing "sufficiently rich non-text edge density". If the intensity of non-text information is high, it means that there is a lot of information in the image that "cannot be expressed by text" (such as charts, diagrams, mixed text and images, object appearance, etc.). In this case, uploading only text will result in the loss of a lot of key information.

[0072] Optional object detection metrics are used to further supplement information about the main objects in the image. If a lightweight object detector is already running on the smart glasses, the following metrics can be added: Target object area ratio This refers to the ratio of the area of ​​the detected target object's bounding box to the total area of ​​the image, i.e. , where bboxj is the bounding box of the j-th detected object, output by the lightweight object detector, area(·) is the calculated pixel area, and I′ is the currently processed image region.

[0073] Object saliency is used to quantify the visual prominence of major objects in an image and their importance to the current interaction task, in order to determine whether it is necessary to retain the non-textual information of these objects by uploading the image. For example, a lightweight object detector on the device can be used to obtain the highest confidence score in the detection results. The higher the score, the more reliable the detected object is and the more likely it is to be the core content that the user intends to focus on. At the same time, the number of detected objects can be counted. The more objects there are, the more complex the image content is and the more likely it contains multiple target information that needs to be retained.

[0074] By introducing the aforementioned multi-dimensional image-side metrics, smart glasses can evaluate the currently collected data from two dimensions: text quality and image quality. In particular, they can accurately identify whether "non-text information is rich and contributes to the task" and "whether there are imaging problems that cause unreliable OCR recognition," providing richer and more accurate basis for refined decision-making on subsequent upload methods.

[0075] After acquiring or calculating multiple secondary indicators to characterize image data quality, the smart glasses fuse these indicators to generate a comprehensive image upload demand score. The design logic of this score is as follows: the better the image data quality, the higher the image upload demand score, indicating a greater preference for uploading images for processing by the target device; conversely, the worse the image data quality, the lower the image upload demand score, indicating a greater preference for uploading locally extracted text data.

[0076] The fusion method can be implemented using various algorithms. One implementation is the weighted summation method, which assigns corresponding weight coefficients to each secondary indicator and calculates the weighted sum as the score for the degree of image transmission required. ,like The weighting coefficients can be configured based on the degree of influence of different indicators on image quality. For example, image sharpness has a greater impact on subsequent recognition results, so it can be assigned a higher weight; image color saturation has a relatively smaller impact, so it can be assigned a lower weight. The weighting coefficients can also be dynamically adjusted according to the current task type. For tasks that require high-precision recognition, the weight of core indicators such as image sharpness can be increased.

[0077] Another approach is the product fusion method, which multiplies the normalized second indicators of each item to obtain a comprehensive score. This method is highly penalizing; if any indicator is too low, it will significantly lower the score for image quality. It is suitable for scenarios where a balance among all image quality indicators is required.

[0078] Step C30: If the required image score is less than the first preset score threshold, then the upload method is determined to be text upload; The smart glasses compare the calculated image transmission score with a pre-set first preset score threshold. The first preset score threshold can be set according to actual application needs, such as an empirical value like 0.6 or 0.7, or it can be configured differently according to different task types.

[0079] When the required image quality score is less than a first preset threshold, it indicates that the overall quality of the current image data is poor, and even if the image is uploaded to the target device, it will be difficult to achieve the desired recognition effect. In this case, if the image is still uploaded, it will not only consume more transmission power, but the poor image quality may also lead to inaccurate recognition results on the target device, making it impossible to effectively execute interactive commands. In contrast, although the locally extracted text data is of medium quality, it still has a certain degree of usability and can provide some useful information to the target device. Therefore, the smart glasses determine that the upload method is to upload text, prioritizing the reduction of transmission power consumption while ensuring basic interactive functions.

[0080] Step C40: If the required image transmission score is greater than or equal to the first preset score threshold, determine the upload method as image upload.

[0081] When the required image quality score is greater than or equal to the first preset score threshold, it indicates that the overall quality of the current image data is good and has good recognizability. In this case, the original image is uploaded to the target device, which is expected to obtain high-quality recognition results and thus accurately execute the interactive commands.

[0082] Through the threshold decision-making mechanism based on the required image quality score, the smart glasses can make refined decisions based on the comprehensive evaluation of image quality when the first indicator is in a moderately blurred area. This mechanism integrates multiple second indicators into a quantifiable decision criterion, avoiding the one-sidedness of single-indicator decisions and achieving a dynamic balance between power consumption optimization and execution accuracy, further improving the adaptability and robustness of the content uploading method.

[0083] In one possible implementation, the step of determining the upload method as image upload if the required image score is greater than or equal to the first preset score threshold includes: Step D10: If the required image transmission score is greater than or equal to the first preset score threshold and less than the second preset score threshold, then the upload method is determined to be uploading a partial image or a compressed image. The smart glasses compare the calculated image transmission requirement score with a first preset score threshold and a second preset score threshold. The second preset score threshold is higher than the first preset score threshold, used to further distinguish the quality level of the image. When the image transmission requirement score is between the first and second preset score thresholds, it indicates that the current image data quality is at a slightly above-average level; it has a certain degree of recognizability, but has not yet reached its optimal state.

[0084] In this scenario, while directly uploading the original image ensures the target device receives complete image information, it consumes significant power during transmission. Furthermore, the image may contain redundant or compressible areas, making it unnecessary to transmit it at its original resolution. Therefore, the smart glasses determine that the upload method should be a partial or compressed image, further optimizing transmission efficiency while ensuring the target device can effectively recognize the text content.

[0085] Uploading a partial image refers to the smart glasses cropping a region from the original image. For example, the smart glasses can determine the bounding box of the recognized text region based on its location coordinates, and appropriately expand the margins outwards by a preset distance (such as expanding by 10% to 20% of pixels) to obtain a partial image. This partial image eliminates the background area from the image, significantly reducing the amount of data transmitted, while retaining all the effective information needed by the target device for text recognition.

[0086] Uploading compressed images refers to the smart glasses performing lossy or lossless compression on the original image or a portion of the image to reduce the image file size. The smart glasses can dynamically adjust the compression rate based on the current image transmission requirement score. The closer the image transmission requirement score is to the second preset score threshold, the better the image quality, and a lower compression rate can be used to retain more image details; the closer the image transmission requirement score is to the first preset score threshold, the more average the image quality, and a higher compression rate can be used to further reduce the amount of data transmitted.

[0087] Furthermore, smart glasses can combine partial image and compressed image methods. First, a partial image of the text area is cropped out, and then this partial image is compressed to maximize the reduction of transmitted data. Through this method, in scenarios with moderate to high image quality, smart glasses avoid the power consumption overhead of uploading the complete original image while providing the target device with sufficient image material for accurate recognition, achieving a further balance between transmission efficiency and recognition accuracy.

[0088] Step D20: If the required image transmission score is greater than or equal to the second preset score threshold, then the upload method is determined to be uploading the original image.

[0089] When the required image quality score is greater than or equal to the second preset score threshold, it indicates that the current image data is of excellent quality and possesses optimal recognition conditions. In this case, the smart glasses determine that the upload method is to upload the original image, that is, to upload the acquired, uncropped or uncompressed original image data to the target device.

[0090] The decision to upload the original image is based on the following considerations: When the image quality reaches an excellent level, the target device can achieve the highest accuracy in text recognition based on the original image, thus maximizing the accuracy of interactive command execution. Although uploading the original image consumes slightly more power than uploading a partial or compressed image, this power consumption is necessary and reasonable considering that the image quality is already optimal, and the complete information contained in the original image may be helpful for advanced processing tasks such as layout analysis and contextual understanding on the target device. Especially for tasks with extremely high accuracy requirements, such as Extract, Key Info, and Verify, uploading the original image can minimize the risk of recognition errors caused by the loss of image information.

[0091] Through the aforementioned multi-level threshold decision-making mechanism based on the required image quality score, the smart glasses further refine the upload method from a binary "upload text or upload image" to three or more levels of options: "upload text," "upload partial or compressed image," and "upload original image." This mechanism enables the smart glasses to adaptively select the most suitable upload method according to different image quality levels. When the image quality is moderate to high, redundant data is reduced through local cropping or compression techniques. When the image quality is excellent, the original image is uploaded to ensure the best recognition effect, achieving a more refined dynamic balance between transmission power consumption and execution accuracy.

[0092] In one possible implementation, the decision rule is a model-based decision, and the step of deciding the upload method based on the decision rule and the first indicator further includes: Step E10: Select a second indicator for evaluating the quality of the image data, wherein the second indicator includes one or more of the following: image sharpness, image contrast, image exposure, image edge density, image grayscale entropy, non-text information intensity, target object area ratio, and target object saliency. Step E20: Input the first indicator and the second indicator into the pre-trained decision model, and output the upload method. The smart glasses use the calculated first and second indicators as input features to the decision-making model, which directly outputs the decision results via upload. This decision-making model is a pre-trained machine learning model, which can be a classification model (such as logistic regression, gradient boosting trees, lightweight neural networks, etc.) used to map the input target indicator to discrete decision categories.

[0093] As an example, the input features of the comprehensive judgment model can consist of three parts, which the smart glasses can tailor according to the computing power of the edge device: Firstly, task type features. When smart glasses receive interaction commands, they parse out the task type, which can include nine categories of task keywords: Extract, Translate, Summarize, Explain, Key Info, Verify, Draft, Share / Save, and Search. Different task types have different requirements for text quality and image information. Introducing task type features allows the model to perceive the specific needs of the current interaction scenario.

[0094] Secondly, text-side OCR features. These are the calculated text quality indicators, i.e., the first indicator. Furthermore, for tasks requiring high accuracy of key fields, such as Key Info or Verify, the confidence scores of key fields can be further extracted, such as the confidence statistics of matched text blocks for dates, amounts, and numbers.

[0095] Third, image-side features. These are the calculated image quality metrics, also known as the second metric. Furthermore, if a lightweight object detector is already running on the edge, detection summary features can be added, such as the number of target objects, the object area ratio, and the highest detection confidence score.

[0096] The decision model can output discrete upload method categories. As an example, the output categories can be refined into five levels, T0 to T4: T0 uploads only OCR text and layout information (such as text line bounding boxes, order, paragraph markers, confidence levels, etc.); T1 uploads structured fields (such as key-value pairs in JSON format); T2 uploads a partially cropped image and the corresponding OCR recognition result; T3 uploads a low-resolution full image and the OCR recognition result; and T4 uploads the original full image. Through multi-level output, the model can select the most refined upload format according to actual needs, minimizing transmission power consumption while ensuring task execution effectiveness.

[0097] The model type can be selected based on the computing resources available on the edge. For example, a logistic regression model can be used, which offers advantages such as extremely low computational cost, ease of interpretation, and stable edge deployment, making it suitable for scenarios requiring significant feature engineering and a limited number of output categories. Alternatively, a gradient boosting tree model (such as LightGBM or XGBoost, trained offline with a tree structure for edge inference) can be employed. This model offers strong fitting capabilities to nonlinear boundaries and is insensitive to feature scale, making it suitable for scenarios with numerous boundary conditions and where reducing the false positive rate is desired. A lightweight multilayer perceptron model can also be used, capable of learning feature interactions, suitable for scenarios requiring stronger generalization capabilities, and supporting INT8 quantization deployment.

[0098] Alternatively, in step E30, the first indicator, the second indicator, and the embedding vector of the image data are input into a pre-trained decision model, and the upload method is output.

[0099] In a more refined implementation, when edge resources allow, smart glasses input not only the target metrics but also the image data embedding vector into the decision model. This embedding vector, a feature representation obtained by inputting the original image into an image encoder, can capture deep semantic information, visual layout, texture features, and other information that is difficult to fully quantify through manually designed metrics. This enhances the ability to discriminate between charts, diagrams, and scenes requiring strong visual retrieval.

[0100] As an example, image representation vectors can be generated by using lightweight convolutional neural networks (such as MobileNet, TinyConv, etc.) to extract the penultimate layer features from the downsampled image, resulting in 32-dimensional, 64-dimensional, or 128-dimensional embedding vectors. Alternatively, color histograms and gradient histograms can be concatenated to form low-dimensional embeddings, achieving even lighter feature extraction. It should be noted that this embedding vector is only used as input to the discriminative model to assist in decision-making and does not replace the actual uploaded image content. The corresponding image upload operation is triggered only when the model determines that an image needs to be uploaded.

[0101] After obtaining the image embedding vector, the smart glasses can concatenate or fuse it with various indicators to form joint features, which are then input into the decision-making model. By introducing the image embedding vector, the model can comprehensively utilize manually designed quality indicators and visual features automatically extracted by deep neural networks to achieve more accurate and robust upload method decisions, which is especially suitable for boundary scenarios such as mixed text and images, complex layouts, and strong visual dependencies.

[0102] It should be noted that the decision-making methods based on machine learning models described above can complement each other, and smart glasses can choose the appropriate method based on factors such as device computing power, power consumption budget, and model size. When computing power is limited, a lightweight model that only inputs the first and second indicators can be used to complete the decision with lower power consumption; when computing power is sufficient and higher accuracy is required, a multimodal model that inputs the first and second indicators and image embedding vectors can be used to obtain richer decision information.

[0103] Through the aforementioned decision-making mechanism based on machine learning models, smart glasses can adaptively learn and optimize the decision-making strategy for uploading methods without the need for manually pre-setting complex threshold rules. This is especially suitable for practical application environments with multiple indicator dimensions, complex decision boundaries, and diverse scene changes, further improving the intelligence level and adaptability of content uploading methods.

[0104] Based on the first and / or second embodiments of the content uploading method of this application, in the third embodiment of the content uploading method of this application, the content that is the same as or similar to the first and second embodiments described above can be referred to the above description and will not be repeated hereafter. Based on this, after the step of acquiring the collected image data, the method further includes: Step F10: If the interactive instruction does not belong to the preset compliant instruction set or does not meet the preset instruction rules, a prompt message is output to remind the user to reissue the instruction. After acquiring image data, the smart glasses first perform compliance verification on the received interactive commands. The preset compliant command set includes all valid command types that the device can recognize and process, such as the aforementioned voice command keywords ("translate," "extract," "save," etc.), specific button combinations, preset head postures, etc. The preset command rules further specify the legality requirements of the commands, such as whether the trigger duration of the command is sufficient, whether the clarity of the voice command meets the recognition requirements, and whether the time interval between consecutive commands is reasonable.

[0105] If the smart glasses determine that the currently received interaction command does not belong to the preset compliant command set, such as the user's voice command "Help me look at this" not matching a corresponding command in the preset keyword library, or the user's head posture movement not reaching the preset trigger amplitude threshold, then the command cannot be correctly recognized by the device. If the interaction command belongs to the compliant command set but does not meet the preset command rules, such as the voice command having severe noise interference leading to low recognition confidence, or the button press duration being insufficient to meet the trigger standard, then the command also cannot be effectively executed.

[0106] As a specific example, when a user issues a "translate" command, but the smart glasses fail to detect any text information during subsequent image acquisition and preliminary text recognition (i.e., the OCR does not recognize any text), although the command itself belongs to the compliant command set, the interaction scenario, combined with the actual recognition result of the image content, fails to meet the execution prerequisites of the command and is therefore considered to have violated the preset command rules. In this case, the smart glasses can output a prompt message through the built-in speaker, vibration motor, or display screen to remind the user to reissue the command. Through this verification mechanism, the smart glasses can effectively filter invalid commands, avoid wasting subsequent processing resources due to command recognition errors or image content not meeting the command requirements, and improve the user interaction experience.

[0107] Step F20: If the interaction instruction belongs to the preset compliant instruction set and satisfies the preset instruction rules, then obtain or determine the task type corresponding to the interaction instruction, and obtain the basic decision rules corresponding to the task type, wherein the basic decision rules include a first rule and a second rule, the first rule is a rule for determining uploaded text, and the second rule is a rule for determining uploaded images; Once the smart glasses determine that the interaction command has passed compliance verification, confirming that the command can be effectively recognized and processed by the device, they then proceed to obtain the preset basic decision rules. These basic decision rules are a set of pre-configured, fast decision-making logics designed to provide an efficient upload method and decision path for scenarios where direct judgment is possible, avoiding the need to enter complex subsequent indicator calculation processes in all cases, thereby further reducing device power consumption and improving response speed.

[0108] The basic decision rules include Rule 1 and Rule 2, and there is no overlap between Rule 1 and Rule 2. Rule 1 is the rule for determining uploaded text, used to identify scenarios where uploaded text should be prioritized regardless of the quality of the text data; Rule 2 is the rule for determining uploaded images, used to identify scenarios where uploaded images should be prioritized regardless of the quality of the text data.

[0109] Step F30: If the first rule is detected to be satisfied, then the upload method is determined to be text upload; The smart glasses match and determine the type of task based on the currently extracted text data or the parsed task type against a first rule. The first rule may include one or more specific scenarios.

[0110] One scenario involves rules based on task type. For example, when the task type is "Translate," the translation task only requires text content and does not require image transmission. The first rule can be configured as "If the task type is 'Translate,' then upload the text directly." Similarly, when the task type is "Draft," since the accuracy requirements for text recognition are relatively low and users typically expect a fast response, the first rule can be configured as "If the task type is 'Draft,' then upload the text directly." Furthermore, when the task type is "Share / Save" and the user's preset preference is "prioritize saving data," the first rule can be configured as "If the task type is 'Share / Save' and the user's preference is data-saving mode, then upload the text directly."

[0111] Another scenario involves rules based on text data. For example, when the text proportion is extremely high and the recognition confidence is extremely high, the first rule can be configured as "if the text proportion is greater than a preset threshold (e.g., greater than 90%) and the average confidence is higher than a preset threshold (e.g., 0.99), then upload the text directly." In this case, even if more complex quality assessments are performed subsequently, it is difficult to obtain a better decision result than directly uploading the text.

[0112] When the text data or task type meets any of the conditions in the first rule, the smart glasses directly determine that the upload method is to upload text.

[0113] Step F40: If the second rule is detected to be satisfied, then the upload method is determined to be uploading an image; The smart glasses match the current text data or task type with the second rule. The second rule can also include one or more specific scenarios.

[0114] One scenario involves rules based on task type. For example, when the task type is Key Info (field extraction) and involves sensitive information, extremely high accuracy is required. A second rule can be configured as "If the task type is Key Info and the extracted field is sensitive information, then upload the image directly." Similarly, when the task type is Verify and the verified content involves legal validity, a second rule can be configured as "If the task type is Verify and the verified content is important evidence, then upload the image directly."

[0115] Another scenario involves rules based on text data. For example, when the number of characters recognized by OCR is extremely small, a second rule can be configured as "If the number of characters recognized by OCR is less than a preset threshold (e.g., less than 5 characters), then upload the image directly." In this case, the amount of text data is too small to reflect complete information, and the risk of recognition errors for a very small number of characters is high. Therefore, uploading the image directly allows the target device to perform more reliable recognition.

[0116] When the text data or task type meets any of the conditions in the second rule, the smart glasses directly determine that the upload method is to upload an image.

[0117] Step F50: If it is detected that neither the first rule nor the second rule is satisfied, then the step of calculating the target index is executed.

[0118] When the smart glasses determine that the current text data and task type do not meet either the first or second rule, it indicates that the upload method cannot be directly determined through the basic decision rules, and a more refined evaluation process is required. At this point, the smart glasses continue to execute subsequent steps, calculate target indicators, and make decisions based on multi-dimensional indicators.

[0119] Through the aforementioned pre-processing mechanism, smart glasses can make rapid decisions in scenarios where judgment is straightforward, avoiding unnecessary indicator calculations and complex decision-making processes, further reducing device power consumption and improving interaction response speed. Simultaneously, this mechanism effectively filters invalid commands through the verification of preset compliant command sets and verifies the prerequisites for command execution based on image content recognition results (e.g., text detection is required for translation commands), avoiding waste of subsequent processing resources. The basic decision-making rules and subsequent refined decision-making mechanisms complement each other, jointly constructing a multi-layered, highly efficient content upload decision-making system from rapid access to refined evaluation, achieving an optimized balance between device power consumption and response efficiency while ensuring the accuracy of interactive command execution.

[0120] In one possible implementation, the step of calculating the target index includes: Step G10: Preprocess the text data to obtain preprocessed text data, wherein the preprocessing includes content merging and / or content filtering; After acquiring image data and performing OCR recognition, smart glasses first preprocess the recognized text data. The purpose of preprocessing is to integrate scattered text information and remove invalid or interfering content, laying the foundation for accurate calculation of subsequent quality indicators.

[0121] Content merging processing aims to address the problem of excessive text fragmentation caused by factors such as text layout and paragraph segmentation during OCR recognition. For example, the bounding box information of each text block can be extracted first, including the left, right, top, and bottom boundaries, and the width, height, and center point coordinates of each text block can be calculated. Then, referring to... Figure 2 As shown, the determination of whether two text blocks should be merged can be based on two spatial relationships: vertical and horizontal relationships and left and right adjacent relationships.

[0122] like Figure 2As shown, for a horizontally overlapping relationship, the width of the horizontally overlapping area of ​​the two text blocks can be calculated and divided by the width of the shorter text block to obtain the horizontal overlap ratio. Simultaneously, the vertical gap between the two text blocks can be calculated and divided by the average height of the two text blocks to obtain the relative vertical distance. When the horizontal overlap ratio is greater than or equal to a preset threshold (e.g., 30%) and the relative vertical distance is less than or equal to a preset threshold (e.g., 0.5 times the line height), the two text blocks can be determined to be in a horizontally overlapping relationship and should be merged.

[0123] like Figure 2 As shown, for left-right adjacent relationships, the horizontal gap between two text blocks can be calculated and divided by the average height to obtain the relative horizontal distance; at the same time, the vertical distance between the center points of the two text blocks can be calculated and divided by the average height to obtain the vertical center difference. When the relative horizontal distance is less than or equal to a preset threshold (e.g., 0.5 times the line height) and the vertical center difference is less than or equal to a preset threshold (e.g., 0.3 times the line height), the two text blocks can be determined to be left-right adjacent and should be merged.

[0124] After identifying all text block pairs that should be merged, a disjoint-set data structure algorithm can be used to group them. Path compression can be used to optimize search efficiency, ultimately resulting in several text block groups. For the text blocks within each group, they can be sorted from top to bottom and from left to right, and the bounding boxes can be merged to form circumscribed rectangles. The text content can then be merged (e.g., connected with spaces).

[0125] Content filtering aims to remove invalid or low-quality information from text data, preventing these distractors from negatively impacting quality assessment results. Three exemplary filtering methods are listed below: Firstly, low-confidence text filtering. All text blocks are traversed, and the OCR confidence value of each block is extracted, with a confidence threshold set (e.g., 0.5). Text blocks with confidence values ​​below this threshold are marked as invalid and filtered out. This strategy effectively filters low-quality recognition results caused by blurry text, noise interference, or recognition errors. For example, in an image containing clear text and a blurred watermark, watermark text with low confidence can be filtered out, while the text content with high confidence is retained.

[0126] Secondly, word count filtering. The system identifies the text block with the most characters and uses its word count as a benchmark, setting a threshold as a preset percentage (e.g., 2%) of the maximum word count. Text blocks with fewer characters than this threshold are filtered out. This strategy effectively removes small, distracting text such as page numbers, watermarks, and button labels, allowing subsequent metric calculations to focus on the main content.

[0127] Third, filtering of small text regions at the edges. The area of ​​each text block is calculated (e.g., using the Shoelace formula to calculate the area of ​​a polygon based on the vertex coordinates of the text box), and it is determined whether the text block is located at the image edge. If the nearest distance from the center point of the text block to the image boundary is less than an edge threshold (e.g., 5% of the image width or height), it is determined to be an edge text block. For text blocks that simultaneously meet the edge location criteria and have an area less than a preset threshold (e.g., 0.5% of the total image area), they are marked as needing filtering and removed. This strategy can further focus on the main content in the central area of ​​the image, filtering out non-core text such as headers, footers, watermarks, and decorative elements.

[0128] It should be noted that if the number of valid text blocks is 0 or lower than the preset minimum number of blocks after the above filtering process, it can be determined that the OCR recognition is insufficient and the subsequent quality indicators cannot be effectively calculated. In this case, the decision path of uploading the image can be directly entered.

[0129] The specific threshold parameters of the above merging and filtering strategies can be configured and adjusted according to the actual application scenario, and this embodiment does not impose specific limitations on them.

[0130] Step G20: Based on the information validity of the preprocessed text data, select at least one first indicator for evaluating the quality of the text data, wherein the first indicator includes one or more of text clarity, text proportion and text concentration. The text sharpness index reflects the recognizability and reliability of text in an image. This index is calculated based on confidence information generated during the OCR recognition process. When recognizing each character or text block, the OCR engine typically outputs a confidence value representing the reliability of the recognition; a higher value indicates a more reliable recognition result. Smart glasses can extract the confidence information corresponding to the valid text blocks retained after preprocessing and calculate the overall sharpness score using statistical methods (such as arithmetic mean, weighted average, median, etc.).

[0131] The text proportion metric reflects the degree to which text areas cover the overall image. This metric is characterized by calculating the proportion of image area occupied by the valid text blocks retained after preprocessing. Specifically, the image dimensions (height and width) are obtained, and the total image area is calculated. Then, for each valid text block, the area enclosed by its bounding box is calculated, such as using the Shoelace formula to calculate the polygon area based on the vertex coordinates of the text box. The areas of all valid text blocks are summed to obtain the total text area. Finally, the total text area is divided by the total image area to obtain the text proportion value (ranging from 0 to 1). For example, in a 1920×1080 pixel image, if the cumulative area of ​​all valid text blocks is 518,400 pixels, then the text proportion is 0.25. This metric can be used to determine the type of image content: when the text proportion is below a preset threshold (e.g., 0.05), it indicates that the image may contain a large amount of non-text information (such as charts, pictures, etc.), and it is more likely to choose to upload the image to retain complete information.

[0132] The text concentration index reflects the spatial distribution of text blocks in an image, used to determine whether there are significant spatial relationships between different text regions. If the text is too scattered in the image, it indicates that different text blocks may have specific spatial layout relationships (such as the correspondence between annotation text and graphic elements in a complex flowchart), or there may be a text-image combination relationship (such as text serving as annotations for different image regions). In this case, uploading only the extracted text data may lose spatial location information, causing the target device to be unable to correctly understand the relationships between the text. Therefore, it is necessary to upload images to preserve layout information.

[0133] Through the above steps, the smart glasses achieve a quantitative assessment of the quality of text information in images, transforming the original text recognition results into structured and comparable quality indicators. This provides a reliable data foundation for subsequent decisions on the upload method based on target indicators, enabling the smart glasses to make adaptive upload decisions that balance transmission efficiency and execution accuracy, based on a full evaluation of local processing performance.

[0134] In one possible implementation, the step of selecting at least one first indicator for evaluating the quality of the text data based on the information validity of the preprocessed text data includes: Step H10: If the first indicator includes text concentration, obtain the center point coordinates of each text block in the text data; The smart glasses first extract the center point coordinates of each text block after preprocessing. The center point coordinates can be obtained by calculating the average of the x and y coordinates of all vertices of the text block's bounding box. That is, the midpoint of the left and right boundaries of the bounding box is taken as the center point x coordinate, and the midpoint of the top and bottom boundaries is taken as the center point y coordinate, thus obtaining the set of center point coordinates for each text block.

[0135] Step H20: Calculate the standard deviation of the center point and the average distance between the center points based on the coordinates of all the center points, and calculate the number of text blocks based on the number of text blocks in the text data, wherein the number of text blocks is negatively correlated with the number of text blocks; First, the standard deviation of the center points. This metric reflects the dispersion of the text block in the horizontal and vertical directions. For example, the standard deviation of the x-coordinates of all center points can be calculated separately. and the standard deviation of the y-coordinate The center point standard deviation is obtained by dividing the center point by the image width W and height H respectively for normalization, and finally taking the average of the two values. ,Right now The larger the value, the more dispersed the text blocks are in the image space; the smaller the value, the more concentrated the text blocks are.

[0136] Secondly, the average distance between center points. This metric reflects the spatial distance characteristics between text blocks. For example, the Euclidean distance between each pair of center points of all text blocks can be calculated. Calculate the average of all distances, then divide by the diagonal length of the image to normalize the result and obtain the average distance to the center point. ,Right now , where n represents the total number of text blocks. The larger the value, the farther apart the text blocks are, and the more dispersed the overall distribution; the smaller the value, the closer the text blocks are, and the more concentrated the overall distribution.

[0137] Third, the number of text blocks factor. This indicator is used to eliminate the influence of the number of text blocks on the dispersion assessment. Generally, when the number of text blocks is large, even if the distribution of each text block is relatively concentrated, the standard deviation of the center point and the mean distance may show higher values ​​due to the increased number of blocks, thus leading to an overestimation of dispersion. To avoid this problem, a text block number factor that is negatively correlated with the number of text blocks can be introduced. For example, this factor can be calculated using logarithmic scaling. ,like: Where K is a preset baseline block count threshold. Logarithmic scaling avoids the problem of excessively high dispersion scores when there are too many blocks, making the dispersion assessment more reasonable.

[0138] Step H30: Determine the text concentration degree based on the standard deviation of the center point, the average distance of the center point, and the text block quantity factor, wherein the text concentration degree is negatively correlated with the standard deviation of the center point, the average distance of the center point, and the text block quantity factor, respectively.

[0139] The smart glasses integrate the three sub-indicators mentioned above to obtain the final text concentration score. Since text concentration reflects the degree of concentration of text distribution, and the standard deviation of the center point, the average distance between the center points, and the number of text blocks are all positively correlated with the degree of dispersion (i.e., the larger the value, the more dispersed), text concentration is negatively correlated with all three sub-indicators.

[0140] For example, a weighted summation method can be used to calculate the concentration of words. ,For example, ,in, , , All of these are preset weighting coefficients with values ​​ranging from 0 to 1.

[0141] This embodiment employs a collaborative evaluation of two dimensions: the standard deviation of the center point and the average distance between the center points. This characterizes the distribution characteristics of text blocks from both the perspectives of dispersion and spatial distance, avoiding the potential bias of a single indicator and making the quantification of dispersion more comprehensive and accurate. Secondly, it introduces a text block quantity factor negatively correlated with the number of text blocks, effectively reducing the overestimation bias caused by an excessive number of text blocks, ensuring that the evaluation results truly reflect the spatial layout of the text rather than being influenced by quantity. Finally, by fusing multi-dimensional features into a single text concentration score, it provides a clear and quantifiable basis for subsequent upload method decisions, avoiding the dilemma of difficult weighting when multiple indicators are used concurrently. This calculation method achieves high-quality quantification of text spatial layout with low computational cost.

[0142] Based on the first, second, and / or third embodiments of the content upload method of this application, in the fourth embodiment of the content upload method of this application, the content that is the same as or similar to the above-described embodiments one, two, and three can be referred to the above description and will not be repeated hereafter. Based on this, the step of uploading the corresponding data to the target device according to the described upload method includes: Step I10: Obtain or determine the task type corresponding to the interaction instruction; When smart glasses receive an interaction command, they can simultaneously parse and determine the task type corresponding to that command.

[0143] Step I20: If the upload method is to upload text, then determine the first upload data based on the task type and upload the first upload data to the target device, wherein the first upload data is the text data and the layout structure of the text data, or the structured fields of the text data; When the decision is to upload text, the smart glasses further select or organize the most suitable data format from the text data as the first data to be uploaded, based on the task type.

[0144] For tasks requiring understanding the spatial relationships of text within the original image to aid semantic understanding or layout reconstruction, the initial uploaded data may include text data and its layout structure. Layout information may include the bounding box coordinates of each text block, the order of text lines, paragraph identifiers, the relative positions of text within the image, reading order, and other layout information. For example, when the task type is Summarize or Explain, preserving layout information that includes paragraph structure and reading order helps the target device more accurately understand the logical hierarchy of the article; when the task type is Extract, layout information helps the target device determine which text belongs to the same area or table, thereby improving the accuracy of extraction.

[0145] For tasks requiring high accuracy of specific fields, such as Key Info (field extraction) or Verify (verification), smart glasses can further process text data into structured field formats as the primary upload data. For example, for images containing invoice information, smart glasses can extract key fields such as the recognized invoice number, invoice date, amount, and tax, and organize them into key-value pairs in JSON format for upload. For images containing ID cards, fields such as name, gender, ethnicity, date of birth, and ID number can be structured and uploaded. By uploading structured fields, the target device does not need to perform text recognition and field parsing again, and can directly perform subsequent operations based on the structured data, significantly improving processing efficiency.

[0146] Step I30: If the upload method is to upload an image, then determine the second upload data based on the task type, and upload the second upload data to the target device. The second upload data is the image data, a partial image of the image data and the text data, or a compressed image of the image data and the text data.

[0147] When the decision is to upload an image, the smart glasses further determine the most suitable data format from the image data as the second upload data, based on the task type.

[0148] For certain task types, the target device needs to obtain complete image information for accurate processing. In this case, the second uploaded data can be the original image data. For example, when the task type is Key Info (field extraction) and involves sensitive information (such as ID card number or bank card number), the original image can be uploaded for the target device to perform high-precision recognition to ensure extraction accuracy. When the task type is Verify and the verification content involves legal validity, uploading the original image can retain complete credential information. When the task type is Explain or Search and the image contains complex charts or objects, the original image can provide the target device with the most comprehensive visual information.

[0149] For other task types, the target device only needs to focus on a specific region in the image to complete the task. In this case, the second uploaded data can be a partial image of the image data plus text data recognized by OCR. The smart glasses can crop out a partial image containing the core content based on the location of the text region recognized by OCR or the main object region identified by the object detector, and then upload it. For example, when the task type is Extract and the image contains multiple independent regions, only the region containing the target text can be cropped; when the task type is Search and a significant object is detected in the image, the region containing that object can be cropped for the target device to perform image retrieval. Uploading partial images reduces the amount of data transmitted while preserving the necessary visual information for the target device to complete the task.

[0150] For tasks requiring image uploads but allowing for some quality loss, the second uploaded data can be a compressed image of the image data plus text data recognized by OCR. The smart glasses can dynamically select an appropriate compression rate and format based on the image quality requirements of the task type. For example, when the task type is Share / Save and the user's image quality requirements are not high, a higher compression rate can be used to reduce the amount of data transmitted; when the task type is Summarize and only the main idea of ​​the text needs to be extracted, the image can be moderately compressed before uploading. For T3 level uploads (low-resolution full images), the smart glasses can downsample the image to a lower resolution before uploading, significantly reducing transmission power consumption while preserving the overall layout.

[0151] By employing a refined data upload determination mechanism based on task type, smart glasses can further optimize transmission efficiency and achieve a dynamic balance between power consumption and accuracy, while ensuring that the target device accurately executes interaction commands.

[0152] In one possible implementation, the local map of the image data includes at least one of the following: Step J10: Crop the image region in the image data where the text clarity is less than or equal to a preset threshold; During OCR recognition, some text areas may have low recognition clarity due to blurriness, uneven lighting, or unique fonts. For these areas with insufficient text clarity, relying solely on locally extracted text data may result in errors. However, cropping these areas and uploading them to the target device, where it utilizes its more powerful recognition capabilities for secondary processing, can potentially yield accurate recognition results.

[0153] As an example, smart glasses can identify text blocks with a resolution below a preset threshold based on a calculated text clarity index, paying particular attention to text blocks containing key lines or key fields. For these low-confidence text regions, the borders can be appropriately expanded around their bounding boxes (e.g., by extending them outwards by 10% to 20%) to ensure that the cropped local image contains complete text context information and avoids missing text edges due to overly tight cropping.

[0154] Step J20: Crop the image region in the image data whose edge density meets a preset condition, wherein the preset condition is a connected region whose edge density is greater than a preset density threshold; Areas with high edge density in an image typically correspond to non-textual information such as charts, graphics, complex layouts, and richly textured content. These areas may contain visual information crucial for the execution of interactive commands, but are difficult to fully express through text recognition.

[0155] As an example, smart glasses can perform connected component analysis on a binarized edge map to identify connected regions in the image whose edge density is higher than a preset threshold. Specifically, all connected components can be extracted, sorted according to their area or edge strength, and the largest connected components (such as the Top-K regions) or the connected components with the highest edge density can be selected as the cropping targets.

[0156] Step J30: Crop the image region containing the preset target object from the image data.

[0157] In some interaction scenarios, user intent may focus on a specific object in an image rather than the text content. For example, when the task type is Search, a user may want to search for similar products using an image search; when the task type is Explain, a user may want to know the name or function of an object in the image.

[0158] As an example, smart glasses can identify preset target objects (such as computers, goods, signs, animals, vehicles, etc.) in images using a lightweight object detector deployed on the device side. The detector outputs the bounding box coordinates and confidence score for each detected object. The smart glasses can then crop the image region corresponding to the bounding box of the target object with the highest confidence score or the target object most relevant to the current task type. If multiple key objects exist, multiple local images can be cropped separately, or the smallest bounding rectangle region containing all key objects can be cropped.

[0159] It should be noted that the above-mentioned methods for acquiring local images can be used individually or in combination. Through these diverse local image cropping strategies, smart glasses can extract the most valuable information regions based on image content features and interactive task requirements, further optimizing transmission efficiency while ensuring the target device accurately executes interactive commands.

[0160] For example, to aid in understanding the technical concept or principle of the content uploading method combined with the first, second, and third embodiments described above, a specific embodiment is now provided. In this specific embodiment, the content uploading method is applied to smart glasses, referring to... Figure 3 As shown, the smart glasses include a task recognition module 10, an image preprocessing module 20, an OCR result analysis module 30, an image index calculation module 40, a rule decision module 50, an upload content generation / packaging module 60, and a communication module 70. The modules work together to execute the content upload process, as detailed below: The task recognition module 10 receives user-input interactive commands or keywords, parses the corresponding task type, and provides input basis for subsequent decision-making based on task dimensions. The image preprocessing module 20 performs preprocessing operations such as downsampling and grayscale conversion on the acquired raw images to reduce subsequent computational complexity and improve the accuracy of index calculation. The preprocessed images are respectively sent to the OCR result analysis module 30 and the image index calculation module 40.

[0161] The OCR result analysis module 30 performs text recognition-related processing on the preprocessed image, including text block merging, content filtering, and index calculation, to obtain OCR result analysis data (including primary indicators such as text clarity, text proportion, and text concentration), and outputs this data to the rule decision module 50. The image index calculation module 40 calculates image-side indicators (i.e., secondary indicators) such as image clarity, grayscale entropy, edge density, and non-text region edge density based on the preprocessed image, and then synchronously outputs the calculated image indicators to the rule decision module 50.

[0162] The rule decision module 50 receives the task type, OCR result analysis data, and image indicators, makes a comprehensive decision, and outputs upload types T0 to T4, specifying the exact format of the content upload. The upload content generation / packaging module 60 generates the corresponding upload content based on the upload type output by the rule decision module and completes the data packaging process. The communication module 70 is responsible for transmitting the packaged upload content to target devices such as mobile phones, completing the final data upload interaction.

[0163] The definitions of each upload type are as follows: T0: Only upload the OCR text and its layout information, including text line / block bounding boxes, reading order, paragraph identifiers, confidence level, etc. T1: Upload structured fields, presented in key-value JSON format, which can be extracted by the smart glasses or mobile phone. T2: Upload a partially cropped image, along with the corresponding OCR recognition or detection results; T3: Upload a low-resolution full image, along with the corresponding OCR recognition or detection results; T4: Upload the original full image.

[0164] Specifically, the rule-based decision-making module can make decisions based on thresholds or on models.

[0165] Threshold-based decision-making processes, such as Figure 4 As shown, it specifically includes: Step S11: Receive interactive instructions and obtain raw image data and interactive instructions.

[0166] Step S12: Perform image preprocessing, OCR text extraction and text region processing, and output interactive instructions, original image and extracted text to decision process 1.

[0167] Step S13: Decision process 1 performs rule judgment to determine whether text or images can be directly transmitted, or whether there is an incorrect interaction intent that requires termination of the process. If the judgment result is "yes", then proceed to step S00: directly determine whether to upload text, upload images, or prompt an error message; if the judgment result is "no", then proceed to step S14.

[0168] Step S14: Text indicator calculation stage, calculate the first indicators such as text clarity, text proportion, and text dispersion.

[0169] Step S15: Next, proceed to decision process 2 to perform text judgment, judging whether the text information can meet the current task requirements based on the above three indicators. If the judgment result is "no", then proceed to step S02: directly output the uploaded image; if the judgment result is "yes", then proceed to step S16.

[0170] Step S16: Image index calculation stage. Calculate the second index, such as image sharpness, edge density, and edge density of non-text areas, and calculate the required image transmission score based on the fusion of each image index.

[0171] Step S17: Decision process 3 performs image judgment. If the required image transmission score is less than the first preset score threshold, then execute step S01: upload text (T0 or T1), and can further refine the decision to T0 or T1 based on the task type; if the required image transmission score is greater than or equal to the first preset score threshold and less than the second preset score threshold, then decide to upload a partial image (T2) or a low-resolution full image (T3), and can further refine the decision based on the task type; if the required image transmission score is greater than or equal to the second preset score threshold, then decide to upload the original image (T4).

[0172] Model-based decision-making processes, such as Figure 5 As shown, it specifically includes: Step S11: Receive interactive instructions and obtain raw image data and interactive instructions.

[0173] Step S12: Perform image preprocessing, OCR text extraction and text region processing, and output interactive instructions, original image and extracted text to decision process 1.

[0174] Step S13: Decision process 1 performs rule judgment to determine whether text or images can be directly transmitted, or whether there is an incorrect interaction intent that requires termination of the process. If the judgment result is "yes", then proceed to step S00: directly determine whether to upload text, upload images, or prompt an error message; if the judgment result is "no", then proceed to step S14.

[0175] Step S14: Text indicator calculation stage, calculate the first indicators such as text clarity, text proportion, and text dispersion.

[0176] Step S16: Image index calculation stage. Calculate the second index, such as image sharpness, edge density, and edge density of non-text areas, and calculate the required image transmission score based on the fusion of each image index.

[0177] Step S18: Proceed to decision process 4, inputting text metrics, image metrics, and original image features into the decision model for multi-dimensional comprehensive evaluation. Based on the evaluation results, decide whether to upload only the OCR text results to save power or upload the original image to ensure accuracy. The decision model can employ lightweight models such as Logistic Regression (LR), Gradient Boosting Decision Tree (GBDT), or Mini Multi-Layer Perceptron (MLP) to complete inference on the edge.

[0178] The closed-loop mechanism of model training, such as Figure 6 As shown, the specific process is as follows: Step S21: Execute online decision-making and interactive operation. Record one sample during each interaction. The sample data includes task type, original OCR output and processed text block, image-side indicators, decision output results and result quality signals (such as whether the cloud requests the original image again, whether the user retakes or cancels, and the confidence level of the cloud model output).

[0179] Step S22: The above interaction samples are uniformly stored in the data record and log library as the basic data source for subsequent model optimization.

[0180] Step S23: Perform interactive tag generation based on log database data. Tags are defined as "minimum sufficient upload type", which is the upload method that saves the most data and power while meeting the quality requirements of the task output. Tags can be constructed through methods such as comparison (running only the OCR route and the image transmission route for the same image simultaneously, and if the outputs are consistent, they tend to be T0 / T1), supplementary transmission signals (if subsequent supplementary transmission is triggered, the tag is adjusted), and manual sampling verification.

[0181] Step S24: Based on the labeled sample data, perform model training and calibration. Lightweight models such as logistic regression, gradient boosting decision trees, or small multilayer perceptrons are used to complete the training and parameter calibration.

[0182] Step S25: Deploy the trained and calibrated model to wearable devices or mobile phones to support online decision-making and interaction processes.

[0183] Step S26: Perform online monitoring and feedback. Monitor the online operation, collect failure samples and update the decision thresholds. Feed the optimized logic back to the online decision-making stage in Step S21 to form a complete closed-loop iteration. The model can be updated offline periodically, with logs summarized regularly to train new models, and then released in a canary manner after regression testing. Alternatively, models or thresholds can be divided according to task type, with a more conservative version used for high-risk tasks, and thresholds continuously calibrated for different scenarios (such as low light, outdoor, screen reflection, etc.).

[0184] It should be noted that the above examples are only used to help understand this embodiment and do not constitute a limitation on the content upload process of this embodiment. Any simple modifications based on this technical concept are within the protection scope of this application.

[0185] Furthermore, embodiments of this application also propose a wearable device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the method described above.

[0186] refer to Figure 7 The diagram illustrates a structural schematic suitable for implementing the embodiments of this application. The wearable devices in the embodiments of this application may also include, but are not limited to, smart glasses, AR headsets, VR headsets, etc. Figure 7 The wearable device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this application.

[0187] like Figure 7 As shown, the wearable device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the wearable device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the wearable device to communicate wirelessly or wiredly with other devices to exchange data. While wearable devices with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0188] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0189] The wearable device provided in this application, employing the content uploading method described in the above embodiments, can solve the technical problems of reducing invalid image transmission processes and lowering data transmission power consumption. Compared with the prior art, the beneficial effects of the wearable device provided in this application are the same as those of the content uploading method provided in the above embodiments, and other technical features of the wearable device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0190] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0191] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0192] In addition, to achieve the above objectives, this application also provides a readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the content uploading method described in the above embodiments.

[0193] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0194] The aforementioned computer-readable storage medium may be included in the wearable device; or it may exist independently and not assembled into the wearable device.

[0195] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by a wearable device, cause the wearable device to implement the process steps of any of the above embodiments.

[0196] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0197] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0198] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the modules themselves.

[0199] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described content uploading method. This solves the technical problems of reducing invalid image transmission processes and lowering data transmission power consumption. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the content uploading method provided in the above embodiments, and will not be repeated here.

[0200] Furthermore, embodiments of this application also propose a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described above.

[0201] The specific implementation method of the computer program product in this application is basically the same as the above-mentioned content uploading method embodiments, and will not be repeated here.

[0202] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0203] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0204] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software sensor. This computer software sensor is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a wearable device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0205] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A content uploading method characterized by comprising: The content upload method includes the following steps: After receiving the interactive command, acquire the collected image data; Text information is extracted from the image data to obtain text data, and at least one first indicator is selected to evaluate the quality of the text data; Obtain or set the decision rule corresponding to the first indicator, and determine the upload method according to the decision rule and the first indicator, wherein the upload method is to upload text or upload an image; The corresponding data is uploaded to the target device according to the upload method described above, so that the target device can execute the interaction command based on the received data.

2. The content uploading method of claim 1, wherein, The step of setting the decision rule corresponding to the first indicator includes: Set an upper and lower threshold for text quality corresponding to the first indicator, to be used for comparison with the first indicator to determine the upload method; or, A decision model is trained based on a preset training dataset to determine the upload method using the trained decision model, wherein the training dataset includes at least the first indicator.

3. The content uploading method of claim 2, wherein, The decision rule is a threshold decision, and the step of determining the upload method based on the decision rule and the first indicator includes: If the first indicator is greater than or equal to the upper limit threshold of text quality, then the upload method is determined to be text upload; If the first indicator is less than or equal to the lower limit threshold of text quality, then the upload method is determined to be uploading an image.

4. The content uploading method of claim 3, wherein, The step of deciding the upload method based on the decision rule and the first indicator further includes: If the first indicator is less than the upper limit threshold of text quality and greater than the lower limit threshold of text quality, then a second indicator for evaluating the quality of the image data is selected, wherein the second indicator includes one or more of the following: image sharpness, image contrast, image exposure, image edge density, image grayscale entropy, non-text information intensity, target object area ratio and target object saliency. The image transmission requirement score is obtained by integrating the second indicators, wherein the higher the quality of the image data, the higher the image transmission requirement score. If the required image score is less than the first preset score threshold, then the upload method is determined to be text upload; If the required image transmission score is greater than or equal to the first preset score threshold, then the upload method is determined to be image upload.

5. The content uploading method as described in claim 4, characterized in that, The step of determining the upload method as image upload if the required image score is greater than or equal to the first preset score threshold includes: If the required image score is greater than or equal to the first preset score threshold and less than the second preset score threshold, then the upload method is determined to be uploading a partial image or a compressed image. If the required image score is greater than or equal to the second preset score threshold, then the upload method is determined to be uploading the original image.

6. The content uploading method as described in claim 2, characterized in that, The decision rule is a model-based decision, and the step of deciding the upload method based on the decision rule and the first indicator further includes: A second indicator is selected for evaluating the quality of the image data, wherein the second indicator includes one or more of the following: image sharpness, image contrast, image exposure, image edge density, image grayscale entropy, non-textual information intensity, target object area ratio, and target object saliency. Input the first indicator and the second indicator into a pre-trained decision model, and output the upload method; or, The first indicator, the second indicator, and the embedding vector of the image data are input into a pre-trained decision model, and the upload method is output.

7. The content uploading method as described in claim 1, characterized in that, The step of selecting at least one first indicator for evaluating the quality of the text data includes: The text data is preprocessed to obtain preprocessed text data, wherein the preprocessing includes content merging and / or content filtering; Based on the information validity of the preprocessed text data, at least one first indicator is selected to evaluate the quality of the text data, wherein the first indicator includes text clarity, text proportion and / or text concentration.

8. The content uploading method as described in claim 7, characterized in that, The step of selecting at least one first indicator for evaluating the quality of the text data based on the preprocessed text data includes: If the first indicator includes text concentration, obtain the center point coordinates of each text block in the text data; The standard deviation of the center point and the average distance between the center points are calculated based on the coordinates of all the center points, and the number of text blocks is calculated based on the number of text blocks in the text data, wherein the number of text blocks is negatively correlated with the number of text blocks; The text concentration is determined based on the standard deviation of the center point, the average distance between the center points, and the number of text blocks, wherein the text concentration is negatively correlated with the standard deviation of the center point, the average distance between the center points, and the number of text blocks, respectively.

9. The content uploading method as described in claim 1, characterized in that, After the step of acquiring the collected image data, the method further includes: If the interactive command does not belong to the preset compliant command set or does not meet the preset command rules, a prompt message will be output to remind the user to reissue the command; If the interaction instruction belongs to the preset compliant instruction set and satisfies the preset instruction rules, then the task type corresponding to the interaction instruction is obtained or determined, and the basic decision rules corresponding to the task type are obtained. The basic decision rules include a first rule and a second rule. The first rule is a rule for determining uploaded text, and the second rule is a rule for determining uploaded images. If the first rule is detected to be satisfied, then the upload method is determined to be text upload; If the second rule is detected to be satisfied, then the upload method is determined to be uploading an image; If it is detected that neither the first rule nor the second rule is satisfied, then the step of selecting at least one first indicator for evaluating the quality of the text data is executed.

10. The content uploading method as described in any one of claims 1 to 9, characterized in that, The step of uploading the corresponding data to the target device according to the upload method includes: Obtain or determine the task type corresponding to the interactive instruction; If the upload method is to upload text, then the first upload data is determined based on the task type, and the first upload data is sent to the target device, wherein the first upload data is the text data and the layout structure of the text data, or the structured fields of the text data; If the upload method is to upload an image, then the second upload data is determined based on the task type, and the second upload data is uploaded to the target device. The second upload data is the image data, a partial image of the image data and the text data, or a compressed image of the image data and the text data.

11. The content uploading method as described in claim 10, characterized in that, The partial map of the image data includes at least one of the following: Cropping image regions in the image data where the text clarity is less than or equal to a preset threshold; Cropping image regions in the image data whose edge density meets a preset condition, wherein the preset condition is a connected region whose edge density is greater than a preset density threshold; The image region containing the preset target object is cropped from the image data.

12. A wearable device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the method as described in any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 12.