javisantana.com

Onpremise

When you start a SaaS company you think every single client will be happy to not have to manage their data, they just forget about servers, backups, maintenance… all those things that you don’t want to deal with.

And that’s actually the reality, you leave those tasks to some people that are experts on that matter and know how to do things far better than you. You only have to pay some amount of money (that of course is less than the one you’d pay if those services were managed within your company) and trust the people behind the service (the hardest part)

At some point in your company you will realize those amazing ideas are not that good when talking about data privacy and security among other things like enterprise authentication methods, custom hardware, redbooth private cloud landing sums up all the reasons quite well.

So yeah, you were using a dozen of amazon services that you can’t use in a private cloud, you will need to face the active directory reality, the high availability, the updates… outside your confy amazon cloud setup, managed by people don’t know anything about your service.

And the problem, there is a little information out there, so you will face a lot of problems, like:

There are some well know companies that have different approaches to the problem:

Do you know more companies doing onpremise/enterprise versions? I’ve created a slack channel so you can join and share them or talk about your experience with the enterprise world problems, feel free to join.

Enterprise Software

Eres un pinpin y te pones a desarrollar software (no necesariamente en este orden), la cosa parece que mola, que no hay que doblar mucho el lomo y es fácil hacer dinero. Te gusta un sector, echas un ojo al software que trata de solucionar la vida de la gente y, como buen ignorante que eres, las primeras frases que salen de tu boca son “JAJAJAJA parece software de los 90”, “esto lo han hecho 4 becarios”, “no tienen ni idea, son unos aficionados”, “esto lo hacemos nosotros mejor con una mano atada a la espalda”. Nota importante: yo he sido el primero en estar en esa situación, así empecé agroguía, de hecho debes estar ahí en algún momento, tienes que ser un veinteañero.

Así que manos a la obra, creas tu proyecto en node.js, TDD, deployment automático, últimas técnicas de desarrollo, el recopón y toda la corporación bendita, después de 2 meses tu software está más o menos listo, ahora solo hay que captar usuarios para tu proyecto B2C.

Pasa 1 año, aprendes que no tenías ni puta idea pero vas malvendiendo, malcaptando clientes/usuarios, te sacan en los periódicos (dándote la falsa sensación de que lo estás haciendo bien) y ganas algún premio de esos que montan por que sobra presupuesto en algún departamente de alguna empresa.

Un día alguien se te acerca, un tío que tiene una empresa cargada de millones, que sabe lo que se hace, joder, que es el puto amo, ve tu software y te plantea: “yo te pagaría XXXX€ por esto si tuviese estas pequeñas modificaciones”. Ese XXXX es 100X de lo que estás cobrando al mes a tus usuarios del SaaS. Llamas a tus socios, lo celebras y sin saberlo estás dentro del mundo llamado “enterprise”, pero aún no eres consciente.

Te comes toda la mierda que el cliente quiere meter en tu producto, llegan otros clientes como ese y rápidamente te das cuenta (bueno, seguramente te lleve meses) que puedes cobrar 100X por lo mismo con mucho menos esfuerzo si empiezas a vender a empresas. Ahora sí, ya sabes que estás en el mundo enterprise, pero aún no sabes cuanta vaselina vas a tener que comprar.

Y aquí es cuando empieza la cosa a ponerse seria, empiezas a hablar con empresas gordas gordas, que tienen gallina de verdad y les dices que tu software es “best in class”. Entonces pasas por diferentes encorbatados y llegas al momento en el que de verdad están interesados. Te mandan la lista de requisitos, la abres pensando que eres el puto amo y entonces es cuando abres el tarro de vaselina.

Y es que además de añadir frases como “big data”, “enterprise ready” y otras perlas que no dicen absolutamente nada pero que debes tener para que los que manejan la gallina sepan posicionarte (sí, eres programador y sabes que todo eso es bullshit, pero no pasa nada, es solo un trámite) debes tener una lista de requisitos en tu software de los que no teías ni idea. Te dejo algunos de ellos.

En resumen, la aplicación de node que parecía simple y ágil ahora tiene una serie de limitaciones, añadidos de clientes, vaya, se ha hecho mayor y ahora miras a esas empresillas que empiezan con una sonrisa (pero por si acaso no dejas de mirarlas), con más pasta en el bolsillo y con suerte habrás dejado de ir a la oficina en bicicleta e irás en un coche como debe ser. Es posible que seas también menos feliz, pero no sabrás si es la edad o tu software enterprise.

como debe ser

Resumen 2015

Este año, igual que el anterior, seré breve. De hecho no tenía pensado escribir este post pero leer los resúmenes tras unos meses (o años) pone las cosas en perspectiva.

A nivel profesional ha sido un año increíble, todo el mundo conoce CartoDB, la gente cree que eres el puto amo, me han conocido en aviones, trenes y bares, he dado charlas, etc, etc. La otra parte, la de trabajar duro, ha sido y está siendo jodida, todo pasa tan rápido que no te da tiempo saborear nada.

A nivel personal ha sido lamentable. Este año que viene será mejor seguro, de hecho empiezo el año mudándome a Madrid después de 4 años y medio con medio pie allí y medio aquí. Por poner una nota positiva, he podido tachar uno de mis TODO vitales, conducir en Nürburgring:

nurg

23 Millones

23 millones

O dicho de otra forma, más de 3000 millones de pesetas (he tenido que revisar dos veces la cifra) es la gallina que unos inversores están arriesgando por un producto “casi español”. Para algo más de información sobre todo este chiringuito leete el post de Miguel Arias que lo explica bastante bien..

La cifra acojona, igual que hacían los 8 millones de la ronda A del año pasado. La mayoría de la gente te felicita, creo que creen que esos millones significan que eres rico de un día para otro y ese dinero sirve parau subirte el sueldo y comprar gilipolleces para la oficina. Nada más lejos de la realidad, esos millones significan libertad, sí, pero muchísima presión, mucho trabajo para gastarlos como se debe. De hecho la celebración cuando nos enteramos fueron unos aquarius que nos tomamos Miguel Arias y yo en el bar de la lado de la oficina, seguramente no nos hubiese entrado nada más, teníamos algún tipo de obstrucción en la garganta. El resto de gente del exec de CartoDB estaba en en diferentes partes del mundo dando el callo…

Más o menos puedes hacerte una idea de lo que significa ese dinero para una empresa como CartoDB, no entraré en detalles, pero a grandes rasgos permite pasar a primera división (puestos de descenso eso sí), puedes pensar a lo grande, cambiar tecnologías que todo el mundo piensa que no puedes cambiar o que son inaccesibles, acceder a personas que de otra forma no podrías. Hace 1 año se acabó el ir en coche por un camino, ahora toca crear la vía para la locomotora.

Pero he venido aquí a hablar de mi libro, qué significa esta historia a nivel personal. El año pasado pasé de ser un desarrollador a ser el CTO. Pasas de preocuparte por tu código a gestionar gente, organizar el trabajo, hacer gestiones que no tienen mucho sentido en ese momento, resolver problemas que no deberían tocarte y en los ratos libres hacer las cosas que realmente tienes que hacer (porque quien sabe qué es lo que tiene que hacer un CTO?) . Ha sido un año realmente difícil, no ha pasado un solo día sin pensar “pero qué cojones estás haciendo y por qué no estás programando?”, siempre tienes la sensación de que no llegas, de que siempre hay un problema más urgente que resolver. Creo que esa sensación de no llegar es algo que pasa cuando trabajas con gente del nivel que tenemos en el equipo.

A medida que pasan los meses te das cuenta que todas esas ideas que tienes de a donde llevar la tecnología no puedes hacerlas solo, necesitas un equipo que sepa lo que se hace en cada área, necesitas tiempo para validar las ideas felices y sobretodo, necesitas creerte y hacer creer que puedes hacerlo. Hace un año pensar que podríamos mandar un parche (y que te lo acepten) a Postgres, que podríamos poner en producción una versión propia de Varnish, aguantar 100k QPS parecía una utopía, ahora son reales. Ahora sabemos como montar un equipo, como diseñar unos procesos de trabajo para escalar, como hacer que la gente se interese por el mundo de los mapas, como tienes que negociar para poder hacer lo que realmente quieres. Ha costado y sigue costando, nadie dijo que fuese fácil.

Así que ahora sigo sin saber qué hace un CTO, y la verdad me da igual, pero tengo claro lo que quiero hacer y me gusta mucho.

Esto es un poco como conducir, al principio te preocupa como usas cada una de los componentes del coche pero luego te das cuenta que lo que realmente hace que lo hagas bien es mirar lejos, lo más lejos posible. Ahora ya no tenemos carretera donde mirar, así que habrá que construirla (con todos estos putos amos)

Local Search

I was talking with some friends about a geo problem they have:

Given a set of elements with a position and a name in a database give me the N closest to a certain point that match some pattern on the name.

So imagine you have openstreetmap database and want to find the first 300 banks and bars closer to Madrid city center (pretty interesting combination I’d say, in Spain ATMs are in the banks).

So in order to test it I loaded Madrid OSM in a postgres database. I just downloaded the data from geofabrik site and imported using osm2pgsql tool, pretty straightforward.

Then I created some indices for the way and amenity columns (using full text search stuff)

-- gist index for the geometry, full text search for the text
create index on planet_osm_point gist(way);
create index on planet_osm_point using to_tsvector('spanish', amenity)

First try, use Nearest Neighbour search

Since postgis 2.0 we have Nearest Neighbour search (thanks to CartoDB which founded it) that allows to use the geospatial index to sort results (read this blogpost in boundless blog), so the first try was to order by distance operator <-> the results from the text filter.

select way, amenity from planet_osm_point where to_tsvector('spanish', amenity) @@ to_tsquery('ba:*') order by way <->'SRID=900913;POINT(-412661.352370664 4926477.3516323)'::geometry limit 301

This takes around 48ms (everything cached). Looking at the explan analyze out of curiosity I realize spatial index wasn’t being used:

Limit  (cost=23947.86..23948.61 rows=301 width=41) (actual time=48.729..48.780 rows=301 loops=1)
   ->  Sort  (cost=23947.86..24035.30 rows=34976 width=41) (actual time=48.727..48.755 rows=301 loops=1)
         Sort Key: ((way <-> '010100002031BF0D00F8DAD368D52F19C1C324815603CB5241'::geometry))
         Sort Method: top-N heapsort  Memory: 48kB
         ->  Bitmap Heap Scan on planet_osm_point  (cost=1363.07..22333.09 rows=34976 width=41) (actual time=12.120..41.008 rows=17382 loops=1)
               Recheck Cond: (to_tsvector('spanish'::regconfig, amenity) @@ to_tsquery('ba:*'::text))
               ->  Bitmap Index Scan on planet_osm_point_to_tsvector_idx1  (cost=0.00..1354.32 rows=34976 width=0) (actual time=10.889..10.889 rows=17382 loops=1)
                     Index Cond: (to_tsvector('spanish'::regconfig, amenity) @@ to_tsquery('ba:*'::text))
 Total runtime: 48.877 ms

I don’t fully understand how the postgres planner works but sounds like it might be using a bitmap and operation both indices. Increasing the limit does not change anything, I though it could change the selectivity. I tried a search by distance:

SELECT way, amenity FROM planet_osm_point 
WHERE
to_tsvector('spanish', amenity) @@ to_tsquery('ba:*') 
AND
st_dwithin(way, 'SRID=900913;POINT(-412661.352370664 4926477.3516323)'::geometry, 500)

and I get the desired BitmapAnd:

 ...
 ->  BitmapAnd  (cost=1422.90..1422.90 rows=42 width=0) (actual time=14.881..14.881 rows=0 loops=1)
         ->  Bitmap Index Scan on planet_osm_point_index  (cost=0.00..68.32 rows=2121 width=0) (actual time=2.229..2.229 rows=6028 loops=1)
               Index Cond: (way && '010300002031BF0D000100000005000000F8DAD368154F19C1C32481560FC95241F8DAD368154F19C1C3248156F7CC5241F8DAD368951019C1C3248156F7CC5241F8DAD368951019C1C32481560FC95241F8DAD368154F19C1C32481560FC95241'::geometry)
         ->  Bitmap Index Scan on planet_osm_point_to_tsvector_idx1  (cost=0.00..1354.32 rows=34976 width=0) (actual time=12.607..12.607 rows=17382 loops=1)
 ...

Second try, use a recursive query

But that’s not the result I was looking for, I need the first 300 bars and banks closer to my location, no matter if they are 30km away

So I though about having a kind of python generator, a query that returns results on demand until I have enough results.

Reading postgres documentation there is a nice trick in CTE documentation, you can do a recursive query without stop condition that stops where the external query reach the limit.

WITH RECURSIVE t(osm_id, way, amenity, distance) AS (
  SELECT osm_id, way, amenity, 4800.0 as distance from planet_osm_point 
        WHERE
    to_tsvector('spanish', amenity) @@ to_tsquery('ba:*')
      AND 
    st_dwithin(way, 'SRID=900913;POINT(-412661.352370664 4926477.3516323)'::geometry, 4800)
  UNION ALL
    -- select only one row from the previous iteration to know the distance
    SELECT p.osm_id, p.way, p.amenity, prev.distance * 2 as distance 
    FROM planet_osm_point p , (SELECT distance FROM t LIMIT 1) prev
       WHERE 
    to_tsvector('spanish', p.amenity) @@ to_tsquery('ba:*')
      AND 
    st_dwithin(p.way, 'SRID=900913;POINT(-412661.352370664 4926477.3516323)'::geometry, prev.distance * 2)
      AND not
    st_dwithin(p.way, 'SRID=900913;POINT(-412661.352370664 4926477.3516323)'::geometry, prev.distance)
),
results as  (
    select *, st_distance(way,'SRID=900913;POINT(-412661.352370664 4926477.3516323)'::geometry) as real_dist FROM t limit 300
)
SELECT * FROM results order by real_dist

The time for this query is around 48ms so no improvement at all. But that may depend on the number of iterations it needs to do until fetch all the results. Starting with 300 meters takes 4 loops since last bar is 2393 meters away form city center.

If it starts the iteration with a bigger radius, like 1500 (2 iterations), the query time is 25ms. If it only needs to do the fist iteration, it’s 18ms which is much better.

So how do we know what would be a good value to start? Hard to say without some density information stats… luckily postgres has pretty good stats about indices and there are good ways to access it: EXPLAIN and _postgis_selectivity

spain=# explain select 1 FROM planet_osm_point where   to_tsvector('spanish', amenity) @@ to_tsquery('ba:*')       ;
                                              QUERY PLAN
-------------------------------------------------------------------------------------------------------
 Bitmap Heap Scan on planet_osm_point  (cost=1363.07..22245.65 rows=34976 width=0)
   Recheck Cond: (to_tsvector('spanish'::regconfig, amenity) @@ to_tsquery('ba:*'::text))
   ->  Bitmap Index Scan on planet_osm_point_to_tsvector_idx1  (cost=0.00..1354.32 rows=34976 width=0)
         Index Cond: (to_tsvector('spanish'::regconfig, amenity) @@ to_tsquery('ba:*'::text))
(4 rows)

Forget everything but rows=34976 that’s the stimation of total number of bars and banks in the whole table (the real count value is 17k but let’s see if it’s good enough)

In order to know how many points there are in a certain are we can use _postgis_selectivity function. It’s kind of hidden, is what actually postgres stats planner use.

select _postgis_selectivity('planet_osm_point', 'way', st_expand('SRID=900913;POINT(-412661.352370664 4926477.3516323)'::geometry, 5000))
 _postgis_selectivity
----------------------
  0.00751143626570702

(you can get the same value using EXPLAIN but I like _postgis_selectivity)

So knowing the selectivity (the part of total rows selected by that bbox), the total number of bars and banks and the total count (1.8M) we can estimate:

total_rows * percentaje_categories * selectivity
1748797 * 0.02 * 0.00751143626570702 ~ 262.0

So in 10km area we have an stimation of 262 points, so to get 300 points:

density = 262/(10000*10000) // points/m^2
area_to_300 = 300/density
R = sqrt(area_to_300/2*PI);
R -> 4268.93996744354m

So using a radius of 4200 meters the query takes ~18ms as it’s doing a single iteration. It’s important to say that calculate the stats is almost free, it takes less than 1ms (except total count that could be precalculated).

Other nice thing about the recursive query is it can be paginated, so you can find 100 results get the last distance and the next time use that distance to start iterating.

Maybe this case is too simple or have too few points but it’s the best approach I found, any idea?