Utiliser les agents IA
Artificial Intelligence (AI) agents can help read and explain code, modify files, execute commands, prepare submission scripts, analyze logs, and assist with debugging. On a shared computing infrastructure, their use requires a clear distinction between the agentic process, the computing resources, and the model service used.
This page presents general technical principles for using AI agents on the Alliance clusters. It does not constitute institutional endorsement of any particular product or vendor.
Important
To date, the Alliance has not formulated a general recommendation for or against the use of AI agents on its clusters. The subject is still under review. Site-specific or regional partner positions must be distinguished from general Alliance policies.
Understanding Where an AI Agent Runs¶
At least two components must be distinguished:
- The agentic client, which runs in the environment where the agent is launched.
- The language model, which can be provided locally or by an external service, depending on the configuration used.
An agent launched with a user's account generally acts with the permissions granted to that account. Therefore, depending on its authorization mode, it can read files, execute commands, or prepare context for the model.
If the agent is launched on a login node, its process runs on that node. If it is launched within an interactive SLURM allocation, its process runs on the allocated compute node. This distinction is important for adhering to high-performance computing best practices.
Where to Run an AI Agent?¶
| Location | Recommended Use | Notes |
|---|---|---|
| Local workstation or virtual machine | Often the simplest option for hosting the agent. | The agent operates outside the cluster, and the user then connects to resources using authorized SSH, MFA, and SLURM mechanisms. |
| Login node | Limit to very light operations. | Prolonged use or costly agentic commands are not suitable for a login node. |
| Compute node in an interactive allocation | Suitable for development, small tests, validation, and interactive debugging. | Use remains subject to local policies and available network connectivity on compute nodes. |
| SLURM batch job | Standard method for reproducible, long, or costly computations. | The agent can help prepare or analyze the job, but the computation is managed by SLURM. |
| Automation node | Site-dependent. | At Calcul Québec, these nodes are intended for deterministic platforms entirely controlled by the user; an AI agent does not fit this definition. |
Using AI Agents with SLURM¶
Compute-intensive workloads must be executed on compute nodes using the scheduler. See Running Jobs and What is a Scheduler?.
For interactive use, first request an allocation, for example:
Then open a shell within the allocation:
Next, verify that the session is indeed within an allocation:
For long or reproducible experiments, prioritize a batch job with sbatch. The agent can help prepare the script, but each #SBATCH directive should be verified by the user before submission.
Session End
Exit the agent and the interactive shell when they are no longer needed, then check that no unnecessary allocations remain active with:
Security, Confidentiality, and Data Protection¶
Sensitive Data¶
The Alliance's general computing resources are not suitable for storing sensitive data. Before using an AI agent, consult Data Protection, Privacy, and Confidentiality as well as the requirements of your institution and project.
The use of an AI agent does not change the classification or applicable obligations for data.
Permissions Between Users¶
An agent launched by another user does not automatically bypass file system permissions. Standard Unix controls still apply.
The situation is different when a user launches an agent themselves: the agent can then access resources that the account can read or modify, to the extent permitted by its authorization mode.
Important Distinction
"another user cannot read my files" and "my own agent can access files that my account can read" are two different issues.
Data Transmitted to an External Service¶
If the model is provided by an external service, some context elements necessary for processing may leave the local infrastructure, depending on the tool's configuration.
Before any use:
- Do not provide secrets, passwords, private keys, tokens, or MFA codes in a prompt.
- Limit the agent to the necessary project directory.
- Avoid exposing files unrelated to the task.
- Examine requested permissions and proposed commands.
- Verify the policies of the provider, institution, and project.
- Do not use the agent with data for which external use conditions are not clearly established.
Permissions and Human Control¶
The user remains responsible for the commands, modifications, and computations they authorize.
A cautious configuration should notably impose the following principles:
- Explain planned actions before execution.
- Do not launch compute-intensive tasks on login nodes.
- Use SLURM for compute workloads.
- Request confirmation before modifying files.
- Request confirmation before installing software.
- Request confirmation before submitting or cancelling jobs.
- Never delete data without explicit approval.
- Do not browse directories unrelated to the project.
- Never display or modify credentials or secrets.
- Summarize actions performed.
SSH, Public Keys, and Multifactor Authentication¶
SSH keys and multifactor authentication may correspond to distinct authentication steps. An accepted public key does not mean that the MFA step should be bypassed or disabled.
See SSH, SSH Keys, Multifactor Authentication, and Automation in the Context of Multifactor Authentication.
When an agent or application encapsulates an SSH connection, it is necessary to verify that the tool can manage an interactive session and the MFA challenge. If a classic SSH connection works but the application's integration fails, the problem may stem from how that application handles SSH interaction or the terminal.
To diagnose the connection from the same environment as the agent:
Never attach a private key, password, MFA code, or API token to a support ticket.
Automation Nodes¶
Rules applicable to automation nodes may vary by site.
At Calcul Québec, the communicated position is that these nodes should be used by deterministic platforms entirely controlled by the user. An AI agent is not considered to meet this definition. Therefore, a persistent deployment of an agent on this type of node is not appropriate at Calcul Québec.
Preferred architectures are rather:
- A local workstation or VM hosting the agent, with connection to the cluster via supported mechanisms.
- An interactive SLURM allocation for interactive tests on a compute node, where policies and connectivity permit.
- SLURM batch jobs for reproducible or costly computations.
Best Practices¶
- Do not use login nodes for compute-intensive tasks or agentic activity.
- Use SLURM for CPU or GPU jobs.
- Limit the agent to the necessary project directory.
- Maintain human control over modifications, installations, deletions, and job submissions.
- Never provide secrets to the agent or in a support ticket.
- Verify applicable data rules before using an external model or service.
- Independently validate the code, scientific results, and SLURM resources suggested by the agent.
- Document tool, library, and environment versions to ensure reproducibility.
- Close interactive allocations when they are no longer needed.
Information to Provide in a Support Ticket¶
When reporting an issue related to an AI agent, provide the following information if possible:
- The name of the cluster.
- The location where the agent is running: local workstation, VM, login node, or compute node.
- The operating system and tool version.
- The SSH connection method used.
- Whether a SLURM allocation is present or not.
- The exact error message, without secrets.
- Relevant output from a read-only diagnostic.
- The observed difference between a classic SSH connection and the connection initiated by the agent, if applicable.