I Have Set

Answered

I Have Set

I have set

export CLEARML_AGENT_SKIP_PYTHON_ENV_INSTALL=true
export CLEARML_AGENT_SKIP_PIP_VENV_INSTALL=true

in my entrypoint.sh (which runs clearml-agent daemon --queue $QUEUES --create-queue --cpu-only --foreground )

but it appears that tasks still take a long time to set up environments. I expected the whole process to be skipped and for the preinstalled python deps in the docker image (which is running this entrypoint script) to be used.

From task pickup to task "run python file" can be several minutes... which is greater than some of the tasks take themselves.

  				
Posted 
	10 months ago

					More  		
  Report
		
					SmallTurkey79
				
					0
					 × 1

Votes Newest

Answers 54

im not running in docker mode though - im running a clearml worker in a docker container (and then multiplying the container)

  				
Posted 
	10 months ago

					More  		
  Report
		
					SmallTurkey79
				
					0
					 × 1

what if the preexisting venv is just the system python? my base image is python:3.10.10 and i just pip install all requirements in that image. Does that not avoid venv still?

it will basically create a new venv inside the container forking the existing preinistalled stuff (i.e. the new venv already has everything the python system has preinstalled)
then it will call "pip install" on all the "installed packages of the Task.
Which should just check everything is there and install nothing

If you set " CLEARML_AGENT_SKIP_PYTHON_ENV_INSTALL=1" it will do checks and just use the existing system python environment as is.

, I can get 50 tasks to run in the same time it takes to run a single one? i cant imagine the apiserver being a noticeable bottleneck.

50 containers on a single machine would be fine if you have enough RAM/CPU, and yes they would run concurrently.
regrading the time itself, again the spinup time of a Task should be negligible.
Pipeline tasks are not meant to be "threads" they are meant as different functions you want to run on different machines,
This means that if your pipeline is just a set of simple functions that require no cpu/gpu or IO, I'm not sure pipeline steps is the right way to go

Does that make sense?

  				
Posted 
	10 months ago

					More  		
  Report
		
					AgitatedDove14
				
					0
					 × 1

oooh thank you, i was hoping for some sort of debugging tips like that. will do.

from a speed-of-clearing-a-queue perspective, is a services-mode queue better or worse than having many workers "always up"?

  				
Posted 
	10 months ago

					More  		
  Report
		
					SmallTurkey79
				
					0
					 × 1

okay that's a similar setup to mine... that's interesting.
much more in line with my expectation.

  				
Posted 
	10 months ago

					More  		
  Report
		
					SmallTurkey79
				
					0
					 × 1

would those containers best be started from something in services mode?

Yes as long as the machine has enough cpu/ram
Notice that the services mode will start a second parallel Task after the first one is done setting up the env, if running with CLEARML_AGENT_SKIP_PYTHON_ENV_INSTALL, with containers that have git/python/clearml-agent preinstalled it should be minimal.

or is it possible to get no-overhead with my approach of worker-inside-docker?

No do not do that, see above explanation on why CLEARML_AGENT_SKIP_PYTHON_ENV_INSTALL does not work in docker venv mode

i designed my tasks as different functions, based mostly on what metrics to report and artifacts that are best cached (and how to best leverage comparisons of tasks). they do require cpu, but not a ton.

just report a single Task as multiple "titles" then each title is it's own step, then inside the "title" they have different seriese

is there a way for me to toggle CLEARML's log level?

Try to set the python master logger base logging level

  				
Posted 
	10 months ago

					More  		
  Report
		
					AgitatedDove14
				
					0
					 × 1

So "Using env ..." take minutes without any output ?

  				
Posted 
	10 months ago

					More  		
  Report
		
					ManiacalLizard2
				
					0
					 × 1

from task pick-up to "git clone" is now ~30s, much better.

This is "spent" calling apt update && update install && pip install clearml-agent
if you have those preinstalled it should be quick

though as far as I understand, the recommendation is still to not run workers-in-docker like this:

if you do not want it to install anything and just use existing venv (leaving the venv as is) and if something is missing then so be it, then yes sure that the way to go

  				
Posted 
	10 months ago

					More  		
  Report
		
					AgitatedDove14
				
					0
					 × 1

oh it's there, before running task.

from task pick-up to "git clone" is now ~30s, much better.

though as far as I understand, the recommendation is still to not run workers-in-docker like this:

export CLEARML_AGENT_SKIP_PYTHON_ENV_INSTALL=1
  export CLEARML_AGENT_SKIP_PIP_VENV_INSTALL=$(which python)

(and fwiw I have this in my entrypoint.sh )

cat <<EOF > ~/clearml.conf
agent {
    vcs_cache {
        enabled: true
    }

    package_manager: {
        type: pip,
        system_site_packages: true,
    }

}
EOF

  				
Posted 
	10 months ago

					More  		
  Report
		
					SmallTurkey79
				
					0
					 × 1

AgitatedDove14 About why we stay on 1.12.2 : None

  				
Posted 
	10 months ago

					More  		
  Report
		
					ManiacalLizard2
				
					0
					 × 1

sometimes I get "lucky" and see something more like what I expect... total experiment time < 1 min (and I have evidence of this happening. logs start-to-finish in sub-minute). But then other times the same task will take 5-10 minutes.

same worker, same queue, just one worker serving it... I am so utterly perplexed by the variation in how long things take. my clearml API server is running on a beefy 32 core machine and not much else is happening right now...

  				
Posted 
	10 months ago

					More  		
  Report
		
					SmallTurkey79
				
					0
					 × 1

minute of silence between first two msgs and then two more mins until a flood of logs. Basically 3 mins total before this task (which does almost nothing - just using it for testing) starts.

  				
Posted 
	10 months ago

					More  		
  Report
		
					SmallTurkey79
				
					0
					 × 1

I think a proper screenshot of the full log with some information redacted is the way to go. Otherwise we are just guessing in the dark

  				
Posted 
	10 months ago

					More  		
  Report
		
					ManiacalLizard2
				
					0
					 × 1

fwiw - i'm starting to wonder if there's a difference between me "resetting the task" vs cloning it.

  				
Posted 
	10 months ago

					More  		
  Report
		
					SmallTurkey79
				
					0
					 × 1

1.12.2 because some bug that make fastai lag 2x
1.8.1rc2 because it fix an annoying git clone bug

  				
Posted 
	10 months ago

					More  		
  Report
		
					ManiacalLizard2
				
					0
					 × 1

hard to see with your croppout here an there ...

  				
Posted 
	10 months ago

					More  		
  Report
		
					ManiacalLizard2
				
					0
					 × 1

i just need to understand what I should be expecting. I thought from putting it into queue in UI to "running my code remotely" (esp with packages preloaded) should be fairly fast turnaround - certainly not three minutes... i'll have to change my whole pipeline design if this is the case)

  				
Posted 
	10 months ago

					More  		
  Report
		
					SmallTurkey79
				
					0
					 × 1

the timestamps were all that mattered in those.

  				
Posted 
	10 months ago

					More  		
  Report
		
					SmallTurkey79
				
					0
					 × 1

apologies - just trying to keep sensitive data out of screenshot

  				
Posted 
	10 months ago

					More  		
  Report
		
					SmallTurkey79
				
					0
					 × 1

i would love some advice on that though - should I be using services mode + docker and some max # of instances to be spinning up multiple tasks instead?

my thinking was to avoid some of the docker overhead. but i did try this approach previously and found that the container limit wasn't exactly respected.

  				
Posted 
	10 months ago

					More  		
  Report
		
					SmallTurkey79
				
					0
					 × 1

but pretty reliably some proportion of tasks still just take a much longer time. 1m - 10m is a variance i'd really like to understand.

  				
Posted 
	10 months ago

					More  		
  Report
		
					SmallTurkey79
				
					0
					 × 1

Please refer to here None
The doc need to be a bit clearer: one require a path and not just true/false

  				
Posted 
	10 months ago

					More  		
  Report
		
					ManiacalLizard2
				
					0
					 × 1

what if the preexisting venv is just the system python ? my base image is python:3.10.10 and i just pip install all requirements in that image . Does that not avoid venv still?

it's good to know that in theory there's a path forward with almost zero overhead . that's what I want .

is it reasonable to expect that with sufficient workers, I can get 50 tasks to run in the same time it takes to run a single one? i cant imagine the apiserver being a noticeable bottleneck .

  				
Posted 
	10 months ago

					More  		
  Report
		
					SmallTurkey79
				
					0
					 × 1

there is almost zero overhead if your docker container alreadyt has everything (including the agent) preinstalled and you set it with CLEARML_AGENT_SKIP_PYTHON_ENV_INSTALL=1
it then should basically just run the code.

  				
Posted 
	10 months ago

					More  		
  Report
		
					AgitatedDove14
				
					0
					 × 1

im not running in docker mode though

hmmm that might be the first issue. it cannot skip venv creation, it can however use a pre-existing venv (but it will change it every time it installs a missing package)
so setting CLEARML_AGENT_SKIP_PYTHON_ENV_INSTALL=1 in non docker mode has no affect

  				
Posted 
	10 months ago

					More  		
  Report
		
					AgitatedDove14
				
					0
					 × 1

Show more results

Write your answer

52K Views

54 Answers

10 months ago