Processing method, intelligent terminal and storage medium

By identifying the target area and configuring the target effect in the smart terminal, the problem of poor user experience caused by the template-based effect in the existing technology is solved, and personalized effect processing is realized, thus improving the user experience.

CN122223152APending Publication Date: 2026-06-16SHENZHEN TRANSSION HLDG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN TRANSSION HLDG CO LTD
Filing Date
2024-12-10
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing special effects template processing methods cannot guarantee consistent results under different shooting environments and content, resulting in a poor user experience.

Method used

By responding to the first instruction, the target area in the preset screen is identified, and the target effects are configured accordingly, including image processing and adjustment based on preset rules, user selection operations, or AI model recognition of the target area and effects.

Benefits of technology

It enables personalized special effects configurations for different shooting environments and content, thus enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122223152A_ABST
    Figure CN122223152A_ABST
Patent Text Reader

Abstract

The application provides a processing method, an intelligent terminal and a storage medium. The processing method comprises the following steps: S1, in response to a first instruction, confirming a target region in a preset picture; and S2, configuring a target special effect for the target region. According to the technical scheme, the target region in the preset picture is confirmed in response to the first instruction, and the target special effect is configured for the target region, so that the target special effect can be configured in a targeted manner, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, specifically to a processing method, a smart terminal, and a storage medium. Background Technology

[0002] As the functions of smart terminals such as mobile phones and tablets continue to improve, they have gradually become one of the commonly used tools in people's daily lives and work.

[0003] In the process of conceiving and implementing this application, the inventors discovered at least the following problems: existing special effects are relatively template-based. Users have different shooting environments and shooting content. If a uniform special effects template is used, the final effect cannot be guaranteed.

[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention

[0005] To address the aforementioned technical issues, this application provides a processing method, a smart terminal, and a storage medium, which specifically configure target effects to enhance the user experience.

[0006] This application provides a processing method, including: S1, in response to the first instruction, confirms the target area in the preset screen; S2, Configure target effects for the target area.

[0007] In one embodiment, determining the target area in the preset screen includes at least one of the following: The area indicated by the first instruction shall be taken as the target area; The preset image is identified based on preset rules to obtain the target area; Obtain the first selection operation and use the area selected by the first selection operation as the target area.

[0008] In one embodiment, the target area in the preset screen includes at least one of the following: The text area in the preset screen; The image area in the preset screen; The character area in the preset image; The scenic area in the preset image; The building area in the preset screen; The face area in the preset image; The preset area in the preset screen.

[0009] In one embodiment, the method further includes determining a target effect, wherein the method for determining the target effect includes at least one of the following: The effect corresponding to the first instruction is taken as the target effect; Get the second selected operation, and take the effect selected by the second selected operation as the target effect; The target effect is obtained or determined based on the features of the target region; The target special effect is obtained or determined based on the content of the preset screen.

[0010] In one embodiment, taking the effect corresponding to the first instruction as the target effect includes at least one of the following: Obtain the text information of the first instruction; The target special effects are generated based on the text information using a preset AI model; Obtain keywords from the text information and determine the target special effects corresponding to the keywords.

[0011] In one embodiment, configuring target effects for the target region includes: Integrate the target effect and the target area; The shooting parameters of the target area are adjusted based on the target special effects. The shape and size of the target area are adjusted based on the target effect. The display of the target area is adjusted based on the target special effects; The target effects are adjusted based on the display of the target area.

[0012] In one embodiment, integrating the target effect and the target region includes at least one of the following: A static pattern corresponding to the target effect is superimposed on the target area; Overlay the animation scene corresponding to the target effect onto the target area; The target area is superimposed onto the static pattern corresponding to the target effect; The target area is superimposed onto the animation scene corresponding to the target effect; Replace the target area with the target effect.

[0013] In one embodiment, the first instruction includes at least one of the following: Voice commands, gesture commands, button commands, handwriting commands, and image recognition commands.

[0014] This application also provides a smart terminal, including: a memory and a processor, wherein the memory stores a processing program, and when the processing program is executed by the processor, it implements the steps of any of the processing methods described above.

[0015] This application also provides a storage medium storing a computer program that, when executed by a processor, implements the steps of any of the processing methods described above.

[0016] The technical solution of this application responds to the first instruction, confirms the target area in the preset screen, and configures target effects for the target area, thereby improving the user experience. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0018] Figure 1 A schematic diagram of the hardware structure of a mobile terminal to implement the various embodiments of this application; Figure 2 A communication network system architecture diagram provided for an embodiment of this application; Figure 3 This is a flowchart illustrating the processing method according to the first embodiment; Figure 4 This is a schematic diagram of the interface according to the processing method shown in the first embodiment; Figure 5 This is a schematic diagram of the specific process of the processing method shown in the second embodiment.

[0019] The realization of the objectives, functional features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and textual descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation

[0020] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0021] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.

[0022] It should be understood that although the terms first, second, third, etc., may be used herein to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this document, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if," as used herein, may be interpreted as "when," "when," or "in response to determination." Furthermore, as used herein, the singular forms "a," "an," and "the" are intended to also include the plural forms unless the context indicates otherwise. It should be further understood that the terms "comprising," "including," indicate the presence of the stated feature, step, operation, element, component, item, kind, and / or group, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, kinds, and / or groups. The terms "or," "and / or," "including at least one of the following," etc., as used in this application, may be interpreted as inclusive, or mean any one or any combination thereof. For example, "including at least one of the following: A, B, C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C." Similarly, "A, B, or C" or "A, B, and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C." Exceptions to this definition only occur when the combination of elements, functions, steps, or operations is inherently mutually exclusive in some way.

[0023] It should be understood that although the steps in the flowcharts of this application's embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0024] Depending on the context, the words “if” or “suppose” as used here can be interpreted as “when” or “in response to determination” or “in response to detection.” Similarly, depending on the context, the phrases “if determination” or “if detection (of the stated condition or event)” can be interpreted as “when determination” or “in response to determination” or “when detection (of the stated condition or event)” or “in response to detection (of the stated condition or event).”

[0025] It should be noted that step designations such as S1 and S2 are used in this document for the purpose of more clearly and concisely describing the corresponding content, and do not constitute a substantial limitation on the order. In specific implementation, those skilled in the art may execute S2 first and then S1, etc., but these should all be within the protection scope of this application.

[0026] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0027] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.

[0028] Smart terminals can be implemented in various forms. For example, the smart terminals described in this application may include smart terminals such as mobile phones, tablets, laptops, PDAs, smartwatches, personal digital assistants (PDAs), portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as fixed terminals such as digital TVs and desktop computers.

[0029] The following description will use a mobile terminal as an example. Those skilled in the art will understand that, apart from elements specifically designed for mobile purposes, the construction according to the embodiments of this application can also be applied to fixed-type terminals.

[0030] Please see Figure 1 This is a schematic diagram of the hardware structure of a mobile terminal implementing various embodiments of this application. The mobile terminal 100 may include: an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (Audio / Video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111, etc. Those skilled in the art will understand that... Figure 1 The mobile terminal structure shown does not constitute a limitation on the mobile terminal. The mobile terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0031] The following is combined Figure 1 A detailed introduction to each component of the mobile terminal: The radio frequency unit 101 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with the processor 110; additionally, it transmits uplink data to the base station. Typically, the radio frequency unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, and a duplexer. Furthermore, the radio frequency unit 101 can also communicate wirelessly with networks and other devices. The aforementioned wireless communications may use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), TDD-LTE (Time Division Duplexing-Long Term Evolution), 5G, and 6G.

[0032] WiFi is a short-range wireless transmission technology. Mobile terminals, through the WiFi module 102, can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 1 WiFi module 102 is shown, but it is understood that it is not a necessary component of a mobile terminal and can be omitted as needed without changing the nature of the invention.

[0033] The audio output unit 103 can convert audio data received by the radio frequency unit 101 or the WiFi module 102 or stored in the memory 109 into audio signals and output them as sound when the mobile terminal 100 is in call signal receiving mode, call mode, recording mode, voice recognition mode, broadcast receiving mode, etc. Furthermore, the audio output unit 103 can also provide audio output related to specific functions performed by the mobile terminal 100 (e.g., call signal receiving sound, message receiving sound, etc.). The audio output unit 103 may include a speaker, a buzzer, etc.

[0034] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The processed image frames can be displayed on the display unit 106. The image frames processed by the GPU 1041 can be stored in the memory 109 (or other storage media) or transmitted via the radio frequency unit 101 or the WiFi module 102. The microphone 1042 can receive sound (audio data) in operating modes such as telephone call mode, recording mode, and voice recognition mode, and can process such sound into audio data. The processed audio (voice) data can be converted into a format that can be transmitted to a mobile communication base station via the radio frequency unit 101 in telephone call mode. The microphone 1042 can implement various types of noise cancellation (or suppression) algorithms to eliminate (or suppress) noise or interference generated during the reception and transmission of audio signals.

[0035] The mobile terminal 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Optionally, the light sensor includes an ambient light sensor and a proximity sensor. Optionally, the ambient light sensor can adjust the brightness of the display panel 1061 according to the ambient light level, and the proximity sensor can turn off the display panel 1061 and / or backlight when the mobile terminal 100 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. Other sensors that may be configured in the phone, such as fingerprint sensors, pressure sensors, iris sensors, molecular sensors, gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.

[0036] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.

[0037] User input unit 107 can be used to receive input numerical or character information, and generate key signal inputs related to user settings and function control of the mobile terminal. Optionally, user input unit 107 may include touch panel 1071 and other input devices 1072. Touch panel 1071, also known as touch screen, can collect touch operations on or near the user (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near touch panel 1071), and drive corresponding connection devices according to a pre-set program. Touch panel 1071 may include two parts: a touch detection device and a touch controller. Optionally, the touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to processor 110, and can receive and execute commands sent by processor 110. In addition, touch panel 1071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 may also include other input devices 1072. Optionally, other input devices 1072 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc., without being specifically limited here.

[0038] Optionally, the touch panel 1071 may cover the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the processor 110 to determine the type of touch event. Subsequently, the processor 110 provides corresponding visual output on the display panel 1061 based on the type of touch event. Although in Figure 1 In this embodiment, the touch panel 1071 and the display panel 1061 are two independent components to realize the input and output functions of the mobile terminal. However, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to realize the input and output functions of the mobile terminal. The specific implementation is not limited here.

[0039] Interface unit 108 serves as an interface through which at least one external device can connect to mobile terminal 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, and so on. Interface unit 108 may be used to receive input (e.g., data, power, etc.) from the external device and transmit the received input to one or more elements within mobile terminal 100, or it may be used to transmit data between mobile terminal 100 and the external device.

[0040] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a program storage area and a data storage area. Optionally, the program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory 109 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0041] The processor 110 is the control center of the mobile terminal. It connects various parts of the mobile terminal via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 109, and by calling data stored in the memory 109, it performs various functions and processes data of the mobile terminal, thereby providing overall monitoring of the mobile terminal. The processor 110 may include one or more processing units; preferably, the processor 110 may integrate an application processor and a modem processor. Optionally, the application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 110.

[0042] The mobile terminal 100 may also include a power supply 111 (such as a battery) that supplies power to various components. Preferably, the power supply 111 can be logically connected to the processor 110 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.

[0043] although Figure 1 As not shown, the mobile terminal 100 may also include a Bluetooth module, etc., which will not be described in detail here.

[0044] To facilitate understanding of the embodiments of this application, the communication network system on which the mobile terminal of this application is based is described below.

[0045] Please see Figure 2 , Figure 2 This application provides a communication network system architecture diagram. The communication network system is an LTE system based on the universal mobile communication technology. The LTE system includes a UE (User Equipment) 201, an E-UTRAN (Evolved UMTS Terrestrial Radio Access Network) 202, an EPC (Evolved Packet Core) 203, and the operator's IP services 204, which are connected in sequence.

[0046] Optionally, UE201 can be the aforementioned terminal 100, which will not be described in detail here.

[0047] E-UTRAN202 includes eNodeB2021 and other eNodeB2022s. Optionally, eNodeB2021 can connect to other eNodeB2022s via backhaul (e.g., X2 interface). eNodeB2021 connects to EPC203 and can provide UE201 with access to EPC203.

[0048] EPC203 may include an MME (Mobility Management Entity) 2031, an HSS (Home Subscriber Server) 2032, other MMEs 2033, an SGW (Serving Gateway) 2034, a PGW (Packet Data Network Gateway) 2035, and a PCRF (Policy and Charging Rules Function) 2036, etc. Optionally, MME2031 is the control node that handles signaling between UE201 and EPC203, providing bearer and connection management. HSS2032 is used to provide registers to manage functions such as the Home Location Register (not shown in the figure) and stores user-specific information such as service characteristics and data rates. All user data can be sent through SGW2034. PGW2035 can provide UE 201 IP address allocation and other functions. PCRF2036 is the policy and charging control decision point for service data flow and IP bearer resources. It selects and provides available policy and charging control decisions for the policy and charging enforcement function unit (not shown in the figure).

[0049] IP services 204 may include the Internet, intranet, IMS (IP Multimedia Subsystem), or other IP services.

[0050] Although the above description uses the LTE system as an example, those skilled in the art should know that this application is not only applicable to the LTE system, but also to other wireless communication systems, such as GSM, CDMA2000, WCDMA, TD-SCDMA, 5G and future new network systems (such as 6G), etc., without limitation.

[0051] Based on the above-described mobile terminal hardware structure and communication network system, various embodiments of this application are proposed.

[0052] First Embodiment Reference Figure 3 , Figure 3 This is a flowchart illustrating the processing method according to the first embodiment. The processing method of this application embodiment can be applied to a smart terminal and includes the following steps: S1: In response to the first command, confirm the target area in the preset screen.

[0053] Optionally, the preset screen may include dynamically generated real-time video streams, images, videos, etc.

[0054] In one embodiment, determining the target area in the preset screen includes at least one of the following: The area indicated by the first instruction shall be taken as the target area; Based on preset rules, a preset image is identified to obtain the target area; Get the first selected operation and use the area selected by the first selected operation as the target area.

[0055] Optionally, preset rules are used to automatically identify and analyze specific content in the image, which can be based on a comprehensive judgment of multiple factors such as image features, shape, color, and texture. Optionally, preset rules can also be set according to different application scenarios to automatically identify target areas of interest in the image. For example, in a security monitoring system, it may be necessary to identify moving objects such as vehicles; in a selfie scenario, it may be necessary to identify human figures or facial features.

[0056] Optionally, the first selection operation includes touch selection, selection on the screen using a mouse, keyboard, or other input device. Optionally, the area selected by the first selection operation is used as the target area, including at least one of the following: detecting an area selected by the user on the screen by dragging with a mouse or other means, and recording the coordinates and range of this area to determine it as the target area; determining an area within a preset range from the click position as the target area based on a detected single click operation; or determining the target area based on the person, object, shape, etc., corresponding to the first selection operation.

[0057] In one embodiment, the target area in the preset screen includes at least one of the following: Preset the text area in the image; Preset the image area in the screen; Preset the area of ​​people in the image; Preset the scenic area in the image; The building area in the preset image; Preset the facial area in the image; The area at the preset position in the preset screen.

[0058] S2: Configure target effects for the target area.

[0059] Optionally, target effects mainly refer to adding effects to photographs in real-time or post-processing through image processing technology or software functions to enhance the visual effect of the photos, express specific emotions, or create a unique artistic style. Optionally, target effects include visual enhancement, atmosphere creation, creative transformation, and dynamic effects. Visual enhancement is used to improve the overall image quality and aesthetics of photos, including HDR (High Dynamic Range) imaging, beautification, and skin smoothing. Atmosphere creation is used to change the tone and atmosphere of photos, conveying different emotions and styles, including retro, black and white, cool-toned, and warm-toned filter effects. Creative transformation adds creativity and fun to photos, including fisheye, wide-angle, shift, and stretching lens effects, as well as adding elements such as halos, rays, and smoke. Dynamic effects incorporate GIF creation, live image shooting, and short video shooting functions, providing more possibilities for target effects.

[0060] In one embodiment, the method further includes determining a target effect, wherein the method for determining the target effect includes at least one of the following: Use the effect corresponding to the first instruction as the target effect; Get the second selected operation, and use the effect selected by the second selected operation as the target effect; Based on the features of the target region, obtain or determine the target special effects; The content of the preset screen is used to obtain or determine the target special effects.

[0061] Optionally, when the first instruction is a voice instruction, the recognition of the first instruction can be based on a preset acoustic neural network model. Optionally, the voice input to the acoustic neural network model is not limited to the voice of a single language, but can include Chinese, English, French, etc.

[0062] Optionally, before recognizing the first instruction based on the acoustic neural network model, the model needs to be trained. Specifically, first, the audio files used to train the acoustic neural network model are determined. These audio files can cover different accents, speech rates, and noisy environments to ensure the model's generalization ability. Simultaneously, corresponding metadata is created for each audio file, including the audio duration and accurate text annotations. Here, the text annotations should be detailed down to each word or even phoneme so that the model learns the precise correspondence between pronunciation and text.

[0063] Secondly, feature extraction from audio files can generally be performed based on filter banks (fbanks) and Mel-frequency cepstral coefficients (MFCCs). Preprocessing steps, such as pre-emphasis, frame segmentation, and windowing, can also be performed during feature extraction to improve feature quality and model training effectiveness.

[0064] Then, an appropriate network architecture is selected to build the acoustic model based on the task requirements. Common network architectures include recurrent neural networks (RNNs), convolutional neural networks (CNNs), long short-term memory networks (LSTMs), Transformers, and Conformers. These network architectures each have their own characteristics; for example, RNNs and LSTMs excel at processing sequential data, CNNs are suitable for capturing local features, while Transformers and Conformers perform exceptionally well in handling long-range dependencies. Depending on the specific application scenario and computing resources, these network architectures can be selected or combined.

[0065] Finally, the acoustic model is trained using a loss function, such as the Connectionist Temporal Classification (CTC) loss function, which can directly map variable-length input sequences to variable-length target sequences without requiring forced alignment of input and output.

[0066] Optionally, during training, the difference between the predicted sequence and the true sequence can be calculated based on the loss function, and the model parameters can be updated using the backpropagation algorithm to minimize this difference. Through multiple iterations of training, the mapping relationship from audio features to text annotations is gradually learned, thereby improving the accuracy of speech recognition.

[0067] Optionally, the instruction information referred to by the first instruction can be directly identified based on an acoustic neural network model.

[0068] Optionally, the user performs a second selection operation on the graphical user interface, including clicking, dragging, or other interactive actions. Optionally, based on the user's selection operation, the type or style of effect the user wants to apply is identified and determined.

[0069] Optionally, the target feature can be determined directly based on the user's second selection operation; alternatively, after determining at least one effect corresponding to the target effect type based on the first instruction, the determined at least one effect can be output and displayed for the user to select. In this way, based on the detection of the second selection operation, the selected effect is taken as the target effect.

[0070] For example, such as Figure 4As shown, if the effect type is determined to be a music effect based on the first instruction, and since more than one music effect is detected, it can be displayed on the interface for the user to select. Correspondingly, after detecting the user's second selection operation, the effect selected by the user is taken as the target effect.

[0071] Optionally, the system analyzes the characteristics of the target area, such as color, shape, and texture, and automatically selects or recommends the most suitable effects for that area. For example, if the target area contains a large amount of green vegetation, effects that enhance green saturation will be recommended.

[0072] Optionally, the overall content of the preset image can be analyzed, including the scene, theme, and mood. Based on the content analysis results, effects that match the image content can be automatically selected or recommended. For example, for a romantic scene, soft focus or highlight effects might be recommended to enhance the atmosphere. Understandably, different content in the target area will correspond to different target effects. In this way, adaptively configuring target effects for the target area makes the application of effects more personalized and precise, meeting the needs of users in different scenarios.

[0073] In one embodiment, the effect corresponding to the first instruction is used as the target effect, including at least one of the following: Obtain the text information of the first instruction; The target special effects are generated based on text information using a pre-set AI model; Extract keywords from the text information and determine the target special effects corresponding to the keywords.

[0074] Optionally, when the first instruction is a voice instruction, it can be converted into text information based on an acoustic neural network model, and then the target effect can be further determined through the text information. Specifically, after converting the first instruction into text information based on the acoustic neural network model, the text information needs to be preprocessed, including removing irrelevant information such as stop words, punctuation marks, and numbers, in order to improve the accuracy and efficiency of text information recognition.

[0075] Optionally, keywords can be extracted from the preprocessed text information using Natural Language Processing (NLP) techniques. Commonly used keyword extraction methods include TF-IDF (Term Frequency-Inverse Document Frequency), TextRank, and LDA (Latent Dirichlet Allocation). Optionally, before obtaining keywords from the text information, a mapping table between at least one keyword and its corresponding effect is pre-defined to determine the effect name or type corresponding to each keyword. Here, the mapping table can be constructed based on manual annotation, machine learning, or expert knowledge.

[0076] Optionally, after obtaining keywords from the text information, the target effect corresponding to the keyword can be determined based on a mapping table. Optionally, if more than two keywords are detected in the text information, determining the target effect may include one of the following: determining more than two effects and selecting the effect most preferred by the user based on historical data as the target effect; simultaneously outputting the two or more determined effects for the user to choose from; automatically selecting the first keyword in the text information and determining the corresponding target effect; or re-identifying the text information.

[0077] In one embodiment, configuring target effects for the target area includes: Integrate target effects and target areas; Adjust the shooting parameters of the target area based on the target effects; Adjust the shape and size of the target area based on the target effect; Adjust the display of the target area based on the target effects; Adjust the target effects based on the display of the target area.

[0078] Optionally, the target effect and target region can be integrated based on the type of target effect to highlight or enhance the target region. Optionally, an image processing neural network model can be pre-built. The network units of the model include, but are not limited to, recurrent neural networks / convolutional neural networks / long short-term memory networks / attention mechanism models / convolutionally enhanced Transformers. Optionally, after building the neural network model, in order to train the neural network model to achieve the application effect of the target effect, the neural network model is trained based on preset or historical data as training data. Here, the training data should include the original image and the effect image after effect processing. The training data should also cover all defined effect data to ensure that the model can learn how to convert ordinary images into images with specific effects. Optionally, choosing an appropriate loss function is crucial when training the neural network model. For example, a cross-entropy loss function can be chosen to measure the difference between the predicted distribution and the true distribution.

[0079] Optionally, shooting parameters include exposure time, aperture size, ISO sensitivity, white balance, etc., which directly affect the visual effects of the image, such as brightness, contrast, and color saturation.

[0080] Optionally, shape and size adjustments typically involve image cropping, scaling, and rotation to alter the image's composition and visual focus, thereby highlighting or weakening specific elements or areas. When applying a target effect, it may be necessary to adjust the shape and size of the target area to suit the effect's presentation or enhance its effect. For example, if a magnifying glass effect is applied, the target area may need to be enlarged; if a rotation effect is applied, the target area may need to be rotated by a corresponding angle. Here, a suitable shape and size adjustment scheme can be automatically calculated based on the requirements of the target effect and the characteristics of the target area, and the image or video content can be processed accordingly.

[0081] Optionally, display adjustments include color correction, contrast enhancement, brightness adjustment, sharpening, and softening. Here, the display style of the target area can be changed accordingly based on the target effect.

[0082] In one embodiment, the target effect and the target area are integrated, including at least one of the following: Overlay a static pattern corresponding to the target effect onto the target area; Overlay the animation scene corresponding to the target effect onto the target area; Overlay the target area onto the static pattern corresponding to the target effect; Overlay the target area onto the animation scene corresponding to the target effect; Replace the target area with the target effect.

[0083] Optionally, when applying a specific effect, it may be necessary to overlay a static pattern related to the effect onto the target area. This pattern can be a texture, icon, border, or preset graphic, used to enhance the visual effect or convey specific information.

[0084] Optionally, the target effect may also require overlaying an animated scene on the target area. This animation could be a looping short video clip, motion graphics, particle effects, etc., to increase the dynamism and appeal of the effect.

[0085] Optionally, for some special effects, it's necessary to overlay the target area itself as a pattern onto another static background or pattern. This operation can change the appearance and visual feel of the target area, making it more consistent with the overall style and atmosphere of the special effect. Specifically, the target area can be extracted from the original image or video and scaled, rotated, cropped, etc., as needed to fit the new background or pattern. Then, the target area is overlaid onto the new background or pattern, and fine-tuned to ensure a harmonious and unified overall effect.

[0086] Optionally, the target area can also be overlaid onto the animation scene. Specifically, based on the characteristics of the animation scene and the needs of the target area, precise overlay and adjustments are made, including dynamic control of parameters such as the target area's transparency, position, and size, to achieve a perfect blend with the animation scene.

[0087] Optionally, depending on the type and requirements of the target effect, the target area can be cropped and pixel adjusted, and at least a portion of the target effect can be replaced by the target area, bringing users a brand-new visual experience.

[0088] In one embodiment, the first instruction includes at least one of the following: Voice commands, gesture commands, button commands, handwriting commands, and image recognition commands.

[0089] In summary, in the processing method provided by the above embodiments, in response to the first instruction, the target area in the preset screen is confirmed; the target area is configured with target effects, which can be configured in a targeted manner to improve the user experience.

[0090] Second Embodiment Reference Figure 5 , Figure 5 This is a schematic flowchart illustrating the processing method according to the second embodiment. Taking the first instruction as a voice instruction as an example, the processing method of this application embodiment includes the following steps: Step S201: The user speaks out the type of photo and effects.

[0091] Optionally, for example, the user can say in English, "capture magic crystal ball." Various languages ​​are supported, such as English, Chinese, and French.

[0092] Step S202: The App or device picks up sound.

[0093] Step S203: Speech recognition and text conversion.

[0094] Optionally, speech recognition can be performed using an acoustic neural network model for speech recognition, converting the acquired speech into text information.

[0095] Step S204: Determine if the keyword matching based on the photo or special effect type is successful. If yes, proceed to step S205; otherwise, return to step S202.

[0096] Optionally, the system detects whether the keyword "taking a photo" is present. For example, if the keyword "capture" is detected, the detection is considered successful. If no text information related to "taking a photo" is detected, the detection is considered unsuccessful. Furthermore, if the action of taking a photo is successfully detected, the system detects target effects, such as "magic crystal ball," which is one type of effect.

[0097] Step S205: Process the image based on the type of special effects and save it.

[0098] Optionally, if the target effect is successfully detected, the camera image is processed, the detected target effect is rendered in real time, and the rendered image is saved.

[0099] In one implementation, if the target effect detection fails, only the action of taking a picture is performed, and the image from the camera is not processed.

[0100] The proposed technical solution allows for targeted configuration of special effects, enhancing the user experience.

[0101] This application also provides a smart terminal, including a memory and a processor. The memory stores a processing program, and when the processing program is executed by the processor, it implements the steps of the processing method in any of the above embodiments.

[0102] This application also provides a storage medium storing a processing program, which, when executed by a processor, implements the steps of the processing method in any of the above embodiments.

[0103] In the embodiments of the smart terminal and storage medium provided in this application, all the technical features of any of the above-described processing method embodiments may be included. The extended and explained contents of the specification are basically the same as the embodiments of the above methods, and will not be repeated here.

[0104] This application also provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to perform the methods described in the various possible implementations above.

[0105] This application also provides a chip, including a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that a device with the chip installed performs the methods described in the various possible implementations above.

[0106] It is understood that the above scenarios are merely examples and do not constitute a limitation on the application scenarios of the technical solutions provided in the embodiments of this application. The technical solutions of this application can also be applied to other scenarios. For example, as those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0107] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0108] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.

[0109] The units in the device of this application embodiment can be merged, divided, and deleted according to actual needs.

[0110] It should be noted that if this application involves the collection or analysis of user behavior data, it will comply with the requirements of relevant laws and regulations, and will be carried out with the user's knowledge and consent in order to protect the user's legitimate rights and interests.

[0111] In this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions are generally described in detail only when they appear for the first time. When they appear again, they are generally not repeated for the sake of brevity. When understanding the technical solutions and other contents of this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions that are not described in detail later can be referred to their previous relevant detailed descriptions.

[0112] In this application, the descriptions of the various embodiments have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0113] The technical features of the present application can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present application.

[0114] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the methods of each embodiment of this application.

[0115] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, storage disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0116] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A processing method, characterized in that, include: S1, in response to the first instruction, confirms the target area in the preset screen; S2, Configure target effects for the target area.

2. The method according to claim 1, characterized in that, Determining the target area in the preset screen includes at least one of the following: The area indicated by the first instruction shall be taken as the target area; The preset image is identified based on preset rules to obtain the target area; Obtain the first selection operation and use the area selected by the first selection operation as the target area.

3. The method according to claim 1, characterized in that, The target area in the preset screen includes at least one of the following: The text area in the preset screen; The image area in the preset screen; The character area in the preset image; The scenic area in the preset image; The building area in the preset screen; The face area in the preset image; The preset area in the preset screen.

4. The method according to any one of claims 1 to 3, characterized in that, It also includes determining the target effect, wherein the method for determining the target effect includes at least one of the following: The effect corresponding to the first instruction is taken as the target effect; Get the second selected operation, and take the effect selected by the second selected operation as the target effect; The target effect is obtained or determined based on the features of the target region; The target special effect is obtained or determined based on the content of the preset screen.

5. The method according to claim 4, characterized in that, The step of using the special effect corresponding to the first instruction as the target special effect includes at least one of the following: Obtain the text information of the first instruction; The target special effects are generated based on the text information using a preset AI model; Obtain keywords from the text information and determine the target special effects corresponding to the keywords.

6. The method according to any one of claims 1 to 3, characterized in that, The configuration of target effects for the target region includes: Integrate the target effect and the target area; The shooting parameters of the target area are adjusted based on the target special effects. The shape and size of the target area are adjusted based on the target effect. The display of the target area is adjusted based on the target special effects; The target effects are adjusted based on the display of the target area.

7. The method according to claim 6, characterized in that, The integration of the target effect and the target region includes at least one of the following: A static pattern corresponding to the target effect is superimposed on the target area; Overlay the animation scene corresponding to the target effect onto the target area; The target area is superimposed onto the static pattern corresponding to the target effect; The target area is superimposed onto the animation scene corresponding to the target effect; Replace the target area with the target effect.

8. The method according to any one of claims 1 to 3, characterized in that, The first instruction includes at least one of the following: Voice commands, gesture commands, button commands, handwriting commands, and image recognition commands.

9. A smart terminal, characterized in that, include: A memory and a processor, wherein a processing program is stored in the memory, and when the processing program is executed by the processor, it implements the steps of the processing method as described in any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium stores a processing program, which, when executed by a processor, implements the steps of the processing method as described in any one of claims 1 to 8.